Lecture 10 Improving Estimators: Rao–Blackwell and Lehmann–Scheffe
In lecture 8 we have seen that a sufficient statistic can be used to summarise all the information in the data about the parameter. In this lecture we will use this idea to construct improved versions of estimators. The Rao–Blackwell theorem states that conditioning an unbiased estimator on the value of never increases its variance. The Lehmann–Scheffe theorem states that, if is a complete sufficient statistic, this procedure gives the unique best unbiased estimator. Finally, Basu’s theorem gives the independence of and for normal samples and this result will be used in lectures 25 and 29. The proofs of these three theorems all use facts about conditional expectations, as shown in propositions A.9 and A.11 in appendix A.
10.1 The Rao–Blackwell Theorem
Assume that we want to estimate a function of the parameter, e.g. we could be interested in the probability of zero counts. As in lecture 2, an estimator is unbiased for , if for all . While unbiased estimators are often easy to find, they often turn out to be wasteful, since they only use a small amount of the available sample. The Rao–Blackwell theorem uses a sufficient statistic to improve an existing estimator by replacing with the conditional expectation , i.e. the average of for all samples which have the same value of as the observed sample.
Let be a sufficient statistic for and let be an unbiased estimator for with for all . Then satisfies the following properties.
-
1.
The random variable does not depend on , but is a function of alone and thus a statistic.
-
2.
The statistic is an unbiased estimator for .
-
3.
For all we have , with equality if and only if .
From definition 8.1 we know that the conditional distribution of the sample given does not depend on . Since is a function of the sample, the conditional distribution and thus the conditional expectation given do not depend on . Thus, is a function of and by definition 1.3 a statistic.
For the second statement, using the tower property from proposition A.9, we find
for all . This shows that is unbiased.
Finally, for the third statement, we can use the law of total variance, proposition A.11, with instead of and instead of . Since we get
Since the conditional variance is non-negative, we find . Equality holds, if and only if and since a non-negative random variable with expectation zero is zero with probability one, this is the case if and only if with probability one, i.e. if and only if coincides with its conditional expectation with probability one. This completes the proof. ∎
Since both estimators are unbiased, the mean squared errors equal the variances, as given in theorem 2.7, and thus is at least as good as , and strictly better unless was already a function of . The procedure of passing from to is called Rao–Blackwellisation. The necessity of sufficiency is highlighted by the fact that conditioning on a statistic which is not sufficient in general will result in a quantity which depends on but is not an estimator.
Let be i.i.d. Poisson with parameter and assume that we want to estimate . The indicator is unbiased for , since . By corollary 8.3 the sum is sufficient. Given , the first observation is distributed as and thus we have
The Rao–Blackwellised estimator now uses all observations. Exercises 10.1 and 10.2 show how to find the conditional distribution and how to apply the same procedure to a continuous model.
10.2 Completeness and the Lehmann–Scheffe Theorem
The Rao–Blackwell theorem improves an existing estimator, but two different unbiased estimators could in principle be improved to two different functions of . The concept of completeness, definition 8.8, solves this problem: if is complete, there can only be one unbiased estimator for among the functions of and thus all Rao–Blackwellisations will converge to this estimator. We start our discussion by naming the target.
An unbiased estimator for with for all is a uniformly minimum variance unbiased estimator (UMVUE) for , if for all and for every unbiased estimator for with finite variance.
The word “uniformly” refers to the parameter: the inequality has to hold for every at once, and there is no reason a priori why such an estimator should exist. The following theorem shows that, whenever a complete sufficient statistic exists, a UMVUE exists as well and is easy to identify.
Let be a complete sufficient statistic for and let be a function of , an unbiased estimator for with for all . Then is a UMVUE for . Furthermore, is the unique UMVUE: any other unbiased estimator for with for all satisfies for all .
To show the uniqueness of as an unbiased estimator for which is a function of , let be another such function. Then we have for all . The function satisfies for all and, since is complete, by definition 8.8, we have for all , i.e. with probability one.
Now let be any unbiased estimator for with finite variance. Then, by the Rao–Blackwell theorem, theorem 10.1, the random variable is a function of which is unbiased for . Thus, by the first step of the proof, it equals with probability one. By the Rao–Blackwell theorem we have then
for all , i.e. is a UMVUE. If for all , then the equality case of the Rao–Blackwell theorem implies with probability one. This completes the proof. ∎
The theorem turns the search for a best unbiased estimator into a two-step recipe: find a complete sufficient statistic , then find any function of which is unbiased for , by adjusting a constant or by Rao–Blackwellising a crude unbiased estimator. For the models of this module the following two facts settle the completeness.
This proposition is given without proof. The first statement rests on the uniqueness theorem for Laplace transforms. The second statement rests on taking derivatives of the relation with respect to . Both statements are allowed to be quoted. Together with corollary 8.3 and exercise 8.2 they form a complete sufficient statistic for all families of the form given in tables 1.1 and 1.2 (with known for the binomial). If we use a one-to-one function of a complete sufficient statistic, the new function is again complete and sufficient. Thus we can use the function instead of if required.
Let be i.i.d. with both parameters unknown. This is a two-parameter exponential family of full rank with natural statistic , and thus the pair is complete and sufficient; exercise 10.4 carries out the verification. Since and are unbiased for and by lemmas 2.2 and 2.3, they are the UMVUEs for these parameters.
The second example puts section 2.3 into perspective: there the biased estimator had smaller mean squared error than . No contradiction, since the Lehmann–Scheffe theorem compares to unbiased estimators only.
10.3 Ancillary Statistics and Basu’s Theorem
A sufficient statistic contains all the information about . At the opposite extreme are statistics which contain none.
A statistic is ancillary for , if the distribution of under does not depend on .
The running example is a normal sample with known variance and unknown mean . Writing we have and the are i.i.d. , for every value of . Thus, the vector of residuals is ancillary for and every function of the residuals is ancillary, e.g. the sample variance and the range ; we will consider the range in exercise 10.5.
One might expect an ancillary statistic to have nothing to do with a sufficient one. Under completeness this becomes a theorem, and it is a surprisingly effective way of proving independence.
Let be a complete sufficient statistic for and let be an ancillary statistic. Then and are independent of each other under for all .
Let be a set of possible values of . Since is ancillary, the probability does not depend on . We compare this probability to the conditional probability
Since is a function of the sample and since is sufficient, the conditional distribution of given does not depend on . Thus, is a function of alone. Using the tower property from proposition A.9 we find
for all . Thus the function satisfies for all . Since is complete, we have with probability one for every .
It remains to show that this implies independence. For any set of possible values of , the indicator function is a function of and can be taken out of the conditional expectation given (see proposition A.10 in appendix A). Using the tower property we find
Since and were arbitrary, we have shown the independence of and and the proof is complete. ∎
The proof uses each hypothesis exactly once: ancillarity makes the unconditional probability constant in , sufficiency makes the conditional probability a statistic, and completeness forces these two to be the same. The most important application of this result in this module is the following.
Let be i.i.d. . Then and are independent.
Fixing the value of , we can consider to be the only parameter. The density of a single observation is
This is of the form given in definition 7.1 with natural statistic and natural parameter . As ranges over , so does , that is, the natural parameter takes all values in , and thus the family is of full rank. By corollary 8.3 and proposition 10.5, the sum , and thus , is a complete and sufficient statistic for . Furthermore, is a function of the residuals and thus, as shown above, is ancillary for . By Basu’s theorem, theorem 10.9, the statistics and are independent under for every and, since was arbitrary, the claim is proved. ∎
This completes the proof of the independence statement of proposition A.3. The result from lectures 25 and 29 on the ratio of to uses this proposition. The proof contains a useful trick: when both parameters are unknown, neither nor can be ancillary and we cannot apply Basu’s theorem directly. The solution to this problem is to fix the value of , i.e. to consider the independence statement for one distribution at a time.
In lecture 13 we will see a second way to derive optimality, using the Cramer–Rao bound, which does not require conditioning.
-
•
Rao–Blackwell: If an unbiased estimator is conditioned on a sufficient statistic , the conditioned estimator is unbiased and has no larger variance.
-
•
A UMVUE is an unbiased estimator which has minimal variance among all unbiased estimators, for all .
-
•
Lehmann–Scheffe: If is complete and sufficient, then the only function of which is unbiased for is the UMVUE for .
-
•
The natural statistic for a full-rank exponential family and the maximum of a uniform sample are complete, but without proof.
-
•
Basu: A complete sufficient statistic is independent of all ancillary statistics. Using this result, one can show that and of a normal sample are independent.
Let be i.i.d. Poisson with parameter , let and let , as in example 10.2. You may use the fact that a sum of independent Poisson random variables is again Poisson, with the parameters added.
-
1.
Show that is an unbiased estimator for and determine the variance of this estimator.
-
2.
Show that for and we have
i.e. that the conditional distribution of given is and that this distribution does not depend on .
-
3.
Deduce that .
-
4.
Show directly that is an unbiased estimator for , using the probability generating function of the Poisson distribution with parameter .
-
5.
Explain why is the UMVUE for and discuss the differences between this UMVUE and the MLE , introduced in exercise 5.7.
Let and let be i.i.d. exponential with rate , i.e. for and let . You may use the fact that the sum of independent random variables of this type has density for .
-
1.
Explain why is an unbiased estimator for the mean and why is sufficient for .
-
2.
Let . Use the independence of and to find the joint density of . Deduce the conditional density of given . The result is
and zero elsewhere. Check that this density does not depend on .
-
3.
Show that and conclude that Rao–Blackwellising leads to the sample mean , again unbiased for .
-
4.
Find a second argument for , based on the symmetry of the sample, without any integration.
Let be i.i.d. uniformly distributed on the set where , and let .
Let be i.i.d. , where both parameters are unknown. This is the situation of example 10.7.
-
1.
Write the density of an observation in the form from definition 7.8 and identify the natural statistic and the natural parameter .
-
2.
Show that the family of densities has full rank, in the sense of definition 7.10, by determining the range of .
-
3.
Deduce that is complete and sufficient and argue that the same holds for the pair .
Let be i.i.d. where is known and is unknown. Furthermore, let be the sample range.
-
1.
Show that is ancillary for .
-
2.
Show that and are independent.
-
3.
Now assume that is known and is unknown. Is ancillary for ? Is a statistic?