Lecture 13 The Cramer–Rao Inequality
In lecture 10 we have found the best unbiased estimators by conditioning on a complete sufficient statistic. In contrast, in this lecture we will discuss a question which could be asked about any unbiased estimator: What are the limitations of unbiased estimators? In the situation of lecture 11, we will see that the variance of an unbiased estimator cannot be smaller than . This is the Cramer–Rao inequality, which gives a precise meaning to the word “optimal” and which we will use in lecture 17 to characterise the maximum likelihood estimator. In this lecture we will prove the inequality, generalise it to functions of the parameter, and then see which estimators achieve the bound. We will conclude by considering the uniform distribution, where the conditions for the bound to hold are violated and where it is possible to beat the bound.
13.1 The Inequality
Throughout this lecture, are i.i.d. with density or probability weights , the parameter space is an open interval, and the conditions of section 11.1 hold. The score and the Fisher information are as in lecture 11; we assume in addition that for all , and we write for the joint density of a value of the sample.
The idea of the proof can be stated in advance: the covariance between an unbiased estimator and the score equals one, for any model, and then the Cauchy–Schwarz inequality implies that small Fisher information implies large variance. We state the result for the covariance as a separate lemma, for use later for functions of the parameter.
Let be a statistic such that for all , and assume that the expectation of can be differentiated under the integral sign, i.e. that
where the sum replaces the integral for discrete models, holds. Then we have
for all .
Since the score has mean zero (lemma 11.1), the covariance is given by the expectation of the product . As in equation (11.2), the score of the sample is the ratio of the joint density and thus the density cancels when we take the expectation:
where the integral is over the support of the distribution, which does not depend on , and where we used the assumption (13.1) in the last step. The final integral can be evaluated as . This completes the proof. ∎
The assumption (13.1) is an additional exchange of differentiation and integration. For exponential families, this relation holds for all statistics with finite variance, by a similar argument from analysis as given in the proof of theorem 7.7, but we will not verify this relation for individual examples.
Since is unbiased, we have and using lemma 13.1 we find that the covariance is the derivative of with respect to , i.e. it equals . Using the Cauchy–Schwarz inequality, lemma A.8, for and , we find
where we used definition 11.2 for the last step. Dividing by the positive number gives the claim. This completes the proof. ∎
Three comments help to read the theorem. First, by proposition 11.4 the bound is : no unbiased estimator can have a standard deviation which decreases faster than the of the sample mean (lemma 2.2). Second, the bound is independent of the estimator used, and only depends on the model: the bound states how much the data can reveal about , regardless of the unbiased method used to estimate it. For the normal mean with known variance the bound is , as shown in example 11.7, and thus achieves this bound. Third, by theorem 2.7, the bound also acts as a lower bound for the mean squared error, but only amongst unbiased estimators: in section 2.3 we have seen a biased estimator for the normal variance, with smaller mean squared error than .
13.2 Estimating a Function of the Parameter
Often the quantity of interest is a function of the parameter, for example the probability of zero counts in example 10.2. If is an unbiased estimator for , the only change in the proof is the value of the covariance.
We have and from lemma 13.1 we know that the covariance satisfies . Using the Cauchy–Schwarz inequality we find and dividing by completes the proof. ∎
Theorem 13.2 is the case ; it is the general form which is worth memorising: identify the function the estimator is unbiased for, differentiate, square and divide by the information. Both numerator and denominator depend on how the model is parametrised, but by proposition 11.10 the bound itself does not: it is a property of the estimator and the model, not of the name we give to the parameter. Exercise 13.6 verifies this.
13.3 Efficiency and Attainment
We give a name to the estimators which attain the bound.
Let be an unbiased estimator for that satisfies the assumptions of theorem 13.3. Then the efficiency of is given by
where the efficiency is the ratio between the Cramer–Rao bound and the actual variance. The estimator is called efficient, if , i.e. if , for all .
From theorem 13.3 we know that the efficiency is contained in the interval , and since the bound is proportional to , an efficiency of means that the estimator uses the information from observations as well as an efficient estimator would use the information from . An efficient estimator is a UMVUE in the sense of definition 10.3, because no unbiased estimator can be better than the bound, but the converse statement is not true, since the bound is not always attainable: in exercise 13.4, the UMVUE for the rate of the exponential distribution has efficiency . These two results complement each other: the Lehmann–Scheffe theorem allows us to identify the best unbiased estimator, whereas the Cramer–Rao bound provides an absolute scale for comparing estimators.
The proof tells us when the bound is attained: equality in the Cramer–Rao inequality is equality in the Cauchy–Schwarz step, and this only holds for random variables which are proportional to each other.
With the same assumptions as in theorem 13.3, the estimator is efficient if and only if there is a function such that
and in this case . In particular, if the model is a one-parameter exponential family as in definition 7.1 and the density is
where is continuously differentiable with for all , and where is the natural statistic, then is an efficient estimator for the mean .
Let and be two random variables with mean zero. Then , and . The proof of lemma A.8 in appendix A considers the quadratic , which is non-negative for all real , and the inequality states that the discriminant is non-positive. Efficiency of implies that the inequality is an equality, i.e. that the discriminant is zero and thus that has a real root . Since is the expectation of the non-negative random variable , we have with probability one, which is (13.2) with . Conversely, if with probability one, we have . This is an equality in the Cauchy–Schwarz inequality and thus implies efficiency. To identify we take the covariance of both sides of (13.2) with . The left-hand side gives by lemma 13.1. The right-hand side gives .
For the exponential family, the proof of proposition 11.9 gives the score of a single observation as , and theorem 7.7 identifies with . Summing over we find , and thus
which is (13.2) for , and . The estimator is unbiased for by construction and has finite variance , and the conditions of section 11.1 hold for the family as noted there. This completes the proof. ∎
The condition (13.2) is restrictive. In an exponential family, an efficient estimator must be an affine function of the natural statistic , and the coefficients of any such affine function cannot depend on (since is fixed while varies over ). Thus, the only functions of which can lead to an efficient estimator are affine functions of the mean of the natural statistic, estimated by . For the Poisson distribution, an exponential family with and , the sample mean is an efficient estimator for . Exercise 13.3 shows that the attainment condition is violated in this case, by showing that is not an affine function of . Thus, the UMVUE from example 10.6 is the best one among the unbiased estimators, but does not achieve the bound.
13.4 When the Conditions Fail
As we have seen, the proof of the inequality in this lecture involves two instances of differentiating under the integral sign, once in lemma 11.1 and once in the assumption (13.1). If the support of the distribution moves as does, both of these exchanges will fail, since an additional term is added to the integral which is not included in the exchange.
Let be i.i.d. with a uniform distribution on and . The support of the distribution depends on and the conditions from section 11.1 are violated. Thus, theorem 13.2 does not apply. We can still see what the formula would be for this model: in example 11.11 the three expressions of theorem 11.3 had three different values, and taking the positive of these three values, , and multiplying by as in proposition 11.4, we find the formal expression and the formal bound .
Now consider the estimators for based on the maximum . The MLE is biased by exercise 2.2, and thus the mean squared error alone does not violate the inequality, since the inequality is a statement about unbiased estimators. The bias-corrected estimator is unbiased and by exercise 10.3 has variance
which is of order instead of . An unbiased estimator with variance smaller than the formal bound in this way is said to be super-efficient. The proof fails at the covariance. The score of the sample on the support is the constant , and thus we have instead of : the two sides of (13.1) are on the left and on the right, and the difference is the contribution of the moving boundary of the support.
There is no contradiction: the inequality has hypotheses, and the uniform model violates these hypotheses. Super-efficiency is a property of models where the parameter is an endpoint of the support, where a single observation near the endpoint can be used to estimate much more accurately than for a regular model. In lecture 16 we will prove the consistency of using a direct argument. In lecture 17 we will see that the limiting distribution of is not normal. In lecture 15 we will use simulation to demonstrate the decay of the variance.
-
•
Under the conditions of section 11.1, all unbiased estimators for satisfy . For we get the bound . The proof uses and the Cauchy–Schwarz inequality.
-
•
The efficiency of an unbiased estimator is the ratio of the bound to the variance, with values in . An efficient estimator is a UMVUE, but the converse is not true.
-
•
The bound is saturated if and only if with probability one. For the one-parameter exponential family, the natural statistic is efficient for the mean, and the only functions of which can be used to construct an efficient estimator are affine functions of the mean.
-
•
The bound depends on regularity conditions. For example, for the uniform distribution on , the support of the distribution depends on the parameter. The unbiased estimator has variance , which is smaller than the formal value . Thus, the estimator is super-efficient.
Let be i.i.d. exponentially distributed with rate . Define for and . Use the fact that the chi-squared distribution from definition A.2 coincides with the distribution from table A.2 and that is distributed.
-
1.
Show that has density for and consequently .
-
2.
Using definition A.2, show that .
-
3.
Using the gamma density of and the change of variables formula from proposition A.19, show that you get the same result as in the previous part.
-
4.
Using the mean and variance of the chi-squared distribution, find and . Check your values by computing the moments of a single observation.
Let be i.i.d. Bernoulli with success probability , and let .
-
1.
Write down the Cramer–Rao bound for , using example 11.5.
-
2.
Show that is an efficient estimator for .
-
3.
Verify condition (13.2) for directly, by writing the score in terms of , and check that the resulting equals .
-
4.
Write down the Cramer–Rao bound for the odds , and explain why no efficient estimator for the odds exists.
-
5.
Show that in fact no unbiased estimator for the odds exists at all.
Let be i.i.d. Poisson with mean and let .
-
1.
Write down the Cramer–Rao bound for , using example 11.6, and show that achieves this bound.
-
2.
Verify condition (13.2) for directly, by writing the score in terms of , and by checking that the resulting equals .
-
3.
Explain why no efficient estimator for exists, and why this does not contradict the fact that is the UMVUE for , using example 10.6.
Let , where , be i.i.d. exponential with rate , let and let be the unbiased version of the MLE from lecture 5.
-
1.
Write down the Cramer–Rao bound for , using example 11.8.
-
2.
In exercise 5.3 we found . Determine the efficiency of and describe how it behaves as .
-
3.
Show that no unbiased estimator for can be efficient.
-
4.
Show that is the UMVUE for .
-
5.
Show that is an efficient estimator for the mean , and explain how this is compatible with the previous results.
Let be i.i.d. . Assume first that is known and is the parameter.
-
1.
Show that and find the Cramer–Rao bound for .
-
2.
Show that is an efficient estimator for , using definition A.2.
-
3.
Now assume that is unknown. From lemma 2.3 we know that the sample variance is an unbiased estimator for . Using proposition A.3, find and compare the result with the bound from the first part of the question. Lecture 14 shows that this bound still holds when is unknown. What can you say about , in light of example 10.7?
Let where is a continuously differentiable bijection with for all . Furthermore, let and be the Fisher information for the sample with model parametrised by and by , respectively.
- 1.
- 2.