Lecture 8 Sufficiency
In example 4.2 we have observed ten coin tosses and got seven heads. The likelihood in this example is . This function depends on the data only via the number of heads, but not on the order of heads and tails, i.e. for learning about we can replace the full record of observations by this single number. The idea of this example is made precise in this lecture: A sufficient statistic contains all information about the parameter. The factorisation theorem helps to find sufficient statistics. The concept of sufficiency forms the basis of improvements of estimators in lecture 10.
In this lecture we write for the data, for an observed value, and
for the joint density or joint probability weights of the sample. The vector argument indicates whether we consider the joint function of the sample, or the function of a single observation. This expression, as a function of for fixed data, is the likelihood from lecture 4.
8.1 Sufficient Statistics
A statistic in the sense of definition 1.3 reduces the data to a smaller summary, and usually some information is lost in the process. The following definition expresses that no information about is lost.
A statistic is sufficient for , if the conditional distribution of given the value does not depend on , for all values of .
If is sufficient, anyone who knows the value of can generate a new data set with the same distribution as the original data set, by sampling from the conditional distribution of given . Since this distribution does not depend on , all information about that we could obtain by studying the data set can now be obtained by studying alone.
For the coin tosses in the introduction, exercise 8.1 shows that the definition is correct: if the number of heads is given as , each of the possible arrangements of heads and tails has conditional probability , independent of the value of . Thus, the positions of the heads do not contain any information about the coin and is sufficient.
Two observations help to place the definition. First, the whole sample is always sufficient, since given the whole sample there is no variability left; sufficiency alone therefore says nothing about how much the data have been reduced, and we return to this in section 8.4. Secondly, a one-to-one function of a sufficient statistic carries the same information and is sufficient as well; for the coin tosses the sample mean is thus sufficient too.
8.2 The Factorisation Theorem
Directly verifying definition 8.1 requires knowledge of the distribution of and involves a conditional probability which, for most models, is tedious to compute. The following theorem replaces the computation with an inspection of the joint density.
The statistic is sufficient for , if and only if there are functions and such that
for all data vectors and all .
The function may depend on the data only through , and the function may not depend on at all: the parameter and the data interact only through . Nothing in the statement requires or to be one-dimensional. We give a complete proof for discrete data and then explain what changes in the continuous case.
Since we have discrete observations, the joint probability weight is the probability of observing the data vector .
Assume first that is sufficient and let be a data vector with . Since the event is contained in the event , we have
whenever . Using the sufficiency of the statistic, the conditional probability on the right does not depend on and thus is a function of the data. The first factor is a function of and . If for the given , both sides in (8.1) equal zero since . For the data vectors with for every , which are never observed, we define . Thus (8.1) holds for all and .
Conversely, assume that (8.1) holds and let be a value of and a parameter value with . Summing the factorisations for all data vectors with , we find
where is the last sum which does not depend on . Since the left-hand side is positive, we have . For data with we get
and for with the conditional probability is zero. Thus, in both cases the value does not depend on and is sufficient. This completes the proof. ∎
For continuous data the event has probability zero and the conditional distribution of given cannot be described as a quotient of probabilities anymore. If the statistic can be extended to a change of variables, i.e. if we can find coordinates such that is a bijection, then the argument presented above can be applied with densities instead of probabilities: The joint density of will factor in a similar way, and after integrating out and dividing we will obtain a conditional density of given which does not depend on . In general, no such completion of the proof is possible, and the proof will have to replace the family of measures by one dominating probability measure, for example a weighted sum of countably many of the , such that the density of is a function of . We will not go into detail here and will use theorem 8.2 for continuous models in analogy to the proof for discrete models, with the understanding that the factorisation, similar to the density, only needs to hold up to sets of probability zero.
The factorisation has a consequence for maximum likelihood. For fixed data the likelihood is , and the factor does not affect where the maximum over is attained: the maximum likelihood estimator, whenever it is unique, is a function of any sufficient statistic. The examples of lecture 5 confirm this: the Bernoulli MLE is , and the uniform MLE is the sample maximum, which exercise 8.2 shows to be sufficient.
8.3 Sufficiency in Exponential Families
For the exponential families from lecture 7, no work is needed to find a sufficient statistic, since the densities are already given in factorised form.
Let be an i.i.d. sample from a one-parameter exponential family as given in definition 7.1 and let be the natural statistic. Then is sufficient for .
The same argument shows that for the -parameter exponential families of lecture 7 the vector of summed natural statistics is sufficient for the parameter vector. It is instructive to see the factorisation directly in the standard models.
Let be i.i.d. Poisson with parameter . Then the joint probability weights are given by
which is of the form (8.1) with and . Thus the sum is sufficient for . Similarly, the sample mean is sufficient. For data sets with the same sum, the likelihood functions are proportional: this was already shown in exercise 4.1 for the log-likelihood, and lecture 9 illustrates this in R.
For the normal distribution , where both parameters are unknown, we can find the sufficient statistic by expanding the squares in the exponent of the density. It transpires that the joint density depends on the data only via the pair . Thus, is a sufficient statistic for . Since , the pair maps to one-to-one and thus is also a sufficient statistic. This is the pair of statistics on which the maximum likelihood estimator from example 5.4 is based. If is known, the factor can be moved into and then alone suffices as a sufficient statistic for . The factorisation is worked out in exercise 8.4.
8.4 Minimal Sufficiency and Completeness
As we have seen in section 8.1, the whole sample is always sufficient. For an i.i.d. sample, the vector of ordered observations is also sufficient, since the joint density does not change when the are interchanged. The aim of this section is to characterise the statistic which allows us to reduce the data as much as possible, while still retaining all the information.
A sufficient statistic is minimal sufficient for , if for every sufficient statistic there is a function such that .
A minimal sufficient statistic can thus be computed from every other sufficient statistic. The definition is awkward to check directly, because it refers to all sufficient statistics at once; the following criterion instead compares the likelihood functions of two data sets.
Let be a statistic, such that every data vector has positive likelihood for some , and such that for all data vectors and we have
Then is minimal sufficient for .
The proof of this theorem is only sketched here. The statement of sufficiency follows from theorem 8.2: For every value of we can choose a fixed data vector such that . Then the ratio , by assumption, does not depend on and can be used as while . To see minimality, let be any sufficient statistic. Then we can write as . If , then does not depend on and thus we have by assumption: the value of is already determined by the value of and is a function of .
The criterion states that contains the likelihood function of the data up to a constant factor. We apply the criterion once, to the Poisson sum. The two-parameter normal case is worked out in exercise 8.4.
Let be i.i.d. Poisson with parameter , and let . For two data vectors and , the joint probability weights of example 8.4 give
If , the power of disappears and the ratio is a constant. If the sums differ by , the ratio contains the factor , which is not constant as ranges over . Thus the ratio is free of if and only if , and is minimal sufficient by theorem 8.6.
For the results of lecture 10 we will need one more property of a statistic, concerning functions of the statistic which are unbiased estimators for zero.
A statistic is complete, if for every function with finite mean under for all , the condition for all implies for all .
This means that the only function of a complete statistic which has expectation zero for all parameter values is the zero function. The condition is violated for statistics which do not reduce the data enough: in exercise 8.6 we will see that the whole Bernoulli sample is sufficient, but not complete, whereas the total is complete, as we will see in lecture 10 (proposition 10.5). The property is important because any two unbiased estimators for the same quantity, which are both functions of a complete statistic , differ by a function of with expectation zero, and by completeness this function must be zero with probability one: for each quantity there is at most one unbiased estimator which is a function of a complete statistic. In lecture 10 we will use this uniqueness to turn the Rao–Blackwell improvement into an optimality result. In the lecture we will also discuss which of the standard statistics are complete.
8.5 The Sample Maximum and Minimum
So far we have considered sufficient statistics which are sums, and we have seen the distributions of these sums in appendix A. For the uniform distribution we will consider the sample maximum, and in order to determine the bias or the mean squared error of the maximum, we need to know the distribution of the maximum. The following result gives the distribution of the maximum and minimum of an i.i.d. sample.
Let be i.i.d. with distribution function and density . Furthermore, let and . Then the random variables and have densities
The maximum is at most , if and only if all observations are at most . Using independence we get
Using the chain rule for differentiation with respect to , we find that the maximum has density . Similarly, the minimum is strictly greater than , if and only if all observations are strictly greater than . Thus we get
Using the chain rule for differentiation we find . This completes the proof. ∎
Let be i.i.d., as in example 1.2. The distribution function is for , and the density is on . By proposition 8.9 the sample maximum has density
and outside this interval. This is the density supplied without proof in exercise 2.2, from which we computed there. The support of the maximum moves with , as the support of a single observation does, and exercise 8.2 shows that the maximum is sufficient for , applying the factorisation theorem with some care about the support.
The minimum plays the corresponding role for models where the density is largest at the left-hand end of the support. In exercise 8.5 we see that the minimum of an exponential sample is itself exponentially distributed, but with rate multiplied by .
-
•
A statistic is sufficient for , if the conditional distribution of the data given does not depend on : the data can be reduced to without losing information about .
-
•
The factorisation theorem states that is sufficient if and only if . Using this result, we can see that a unique MLE is always a function of any sufficient statistic.
-
•
For exponential families, the summed natural statistic is sufficient. This includes the Bernoulli total, the Poisson sum and the normal pair .
-
•
A minimal sufficient statistic is a function of any sufficient statistic, and satisfies the condition if and only if the likelihood ratio does not depend on . A complete statistic is any statistic such that the only function of the statistic which has expectation zero for all is the zero function. In lecture 10 we will use this property to identify the unique best unbiased estimator.
-
•
The densities of the sample maximum and minimum are and , obtained by differentiating and , respectively.
Let be i.i.d. Bernoulli with success probability and let be the number of successes. From the remark about sums in appendix A we know that .
Let be i.i.d. with .
-
1.
Using the indicator notation for the density of a single observation, show that the joint density can be written as
-
2.
Deduce from theorem 8.2 that the sample maximum is sufficient for . Identify the functions and .
-
3.
Why can the factor with be included in , but the factor with cannot?
Let be i.i.d. exponential with rate , i.e. for .
Let be i.i.d.
-
1.
Assume that is known and is unknown. Show that is sufficient for .
-
2.
Assume that both parameters are unknown. By expanding the squares in the exponent of the joint density, show that the joint density depends on the data only via the pair . Using theorem 8.2, show that is sufficient for .
-
3.
Using theorem 8.6, show that the pair from the previous part is minimal sufficient for .
-
4.
Is the sample mean on its own sufficient for when both parameters are unknown? Justify your answer using the previous part.
Let be i.i.d. exponential with rate , and let .
-
1.
Show that is exponentially distributed with rate , first by proposition 8.9 and then directly from .
-
2.
Show that is an unbiased estimator for the mean , and compute its variance.
-
3.
Compare the variance of with that of the sample mean , which is also unbiased for , and comment on the behaviour of the two estimators as grows. Is a function of the sufficient statistic ?
Let be i.i.d. Bernoulli with success probability , where , and let be the whole sample.
-
1.
Show that satisfies for all and .
-
2.
Using definition 8.8, show that is not complete, despite being sufficient.