Lecture 17 Asymptotic Normality and Efficiency of the MLE
In lecture 16 we have seen that under regularity conditions the maximum likelihood estimator converges to the true parameter value. In this lecture we will discuss the error in more detail: It transpires that for large sample size the error is approximately normally distributed with mean zero and variance given by the Cramer–Rao bound from lecture 13. Thus, the MLE is unbiased in the limit, and more importantly, is as accurate as any unbiased estimator can be. The proof of this result is a challenging one, but it is based on tools we have already learned in this module. The argument is based on Taylor expansion of the score function, and the resulting theorem allows us to compute standard errors and to derive approximate confidence statements. We will see the resulting numbers in lecture 18 and will make the statements precise in lecture 25.
17.1 The Theorem
Using the notation from lecture 16, we write for the Fisher information of a single observation and for the Fisher information of the sample, as given in definition 11.2 and proposition 11.4. The result is formulated in terms of the scaled error , as this is the only error term which has a non-degenerate limit: by consistency the error goes to zero and the factor is required to balance this effect.
Suppose that the regularity conditions (R1) to (R4) from definition 16.1 are satisfied, that has finite mean under for all , and that for every sample size and every sample where all coordinates are in the support of the model, the likelihood equation has a unique solution which coincides with the maximum likelihood estimator. Then
Before we give the proof, we unpack the statement of the theorem. The theorem informally states that, as gets large,
This statement contains three different aspects: (1) The centre of the approximating distribution equals the true value, i.e. the MLE is asymptotically unbiased (this does not mean that the MLE is unbiased for finite , only that it is asymptotically unbiased). (2) The spread of the distribution decreases as , as for the sample mean. (3) The variance of the approximating distribution equals . This is the Cramer–Rao lower bound from lecture 13: no unbiased estimator can have smaller variance for any and the MLE achieves this bound in the limit. The last property has a name.
An estimator for , which does not have to be the maximum likelihood estimator, is asymptotically efficient, if for all .
In this language theorem 17.1 states that under the regularity conditions the MLE is asymptotically efficient. In lecture 13 we have seen that an efficient estimator, one which achieves the Cramer–Rao bound exactly, only exists for special models and parametrisations, namely the natural statistic of an exponential family. The asymptotic statement is much wider in applicability, as it holds for every regular model and every parametrisation, since by theorem 5.6 the MLE transforms together with the parameter.
17.2 Proof of the Theorem
We will use a Taylor expansion of the score function around the true value. Since the MLE is a root of the score, we can use Taylor expansion of around the point to write the error term in terms of and , which are both sums of i.i.d. terms and thus can be analysed using the limit theorems for sample means.
Let the true value be . Then, by (R3), the function is two times continuously differentiable. Thus, by Taylor’s theorem with the Lagrange form of the remainder, for around the point we get
where is a point between and . Since solves the likelihood equation, the left-hand side is zero. Solving for and multiplying by we can write the scaled error as a ratio:
where
We now take the limit of each of the three terms separately.
The numerator is a scaled sum of i.i.d. terms: by definition of the score we have and using the identities (16.1) we find that each term has mean zero and finite, positive variance . Thus, by the central limit theorem, theorem A.15, for the sample mean of the we get
The first term in the denominator is also a sample mean: and by the law of large numbers, theorem A.14, and the second identity in (16.1) we find
The remainder term is where the consistency from lecture 16 and condition (R4) come into play. Let and be as in (R4). On the event , the intermediate point also lies within of . Thus, writing as a sum of the observations and using (R4) for each term, we find
on this event. By theorem 16.3, the probability of the event goes to one, and by the law of large numbers the right-hand side converges in probability to the finite constant . Thus, for all sufficiently large we have . The other factor of is , which converges in probability to zero by theorem 16.3. For any , the event is contained in the union of the events and , both of which have probability that goes to zero, and thus .
It remains to combine these three limits, and the appropriate tool for this is Slutsky’s theorem, theorem A.16. Since and , the denominator satisfies , where the value is constant and non-zero. By the quotient part of Slutsky’s theorem, for the representation (17.1), we have
where in the last step we used the fact that dividing a random variable by the constant gives a random variable, in this case with and . This completes the proof. ∎
The structure of this argument is typical of most large-sample results in statistics: Taylor expansion is used to transform a question about an implicitly defined estimator into a question about sums of i.i.d. terms, the central limit theorem is used to analyse the leading term, the law of large numbers is used to analyse the coefficient of the leading term, the remainder is shown to be negligible using consistency of the estimator, and Slutsky’s theorem is used to combine the individual results. The only tricky step is the analysis of the remainder term, and condition (R4) is the reason that this term is straightforward to analyse.
17.3 Standard Errors
The variance in theorem 17.1 depends on the unknown , and so cannot be reported as it stands. In practice we replace by its estimate, and the next result says that this does not disturb the limit.
With the same assumptions as in theorem 17.1, we have
By (R3) the function is continuous and positive on and thus the function is continuous at . By from theorem 16.3, the continuous mapping theorem from theorem A.17 implies that . By multiplying the result of theorem 17.1 by this factor and using the product part of Slutsky’s theorem, theorem A.16, we find
and the claim is proved. ∎
The quantity
is called the standard error of the MLE. As a statistic, this quantity from the corollary is an approximation to the standard deviation of for large . The corollary also allows us to justify approximate confidence statements: Since a standard normal variable takes values in with probability , the random interval covers with probability which tends to . Intervals of this form were used in an informal way in lecture 12 and we will make these intervals precise in lecture 25. In lecture 18 we will use simulation to check how close the real coverage of these intervals is to for moderate . If we do not have a closed formula for , the denominator can be replaced by the observed curvature of the log-likelihood at the maximum. This approach is taken by a numerical optimiser and we will use the resulting values in lecture 18.
17.4 Examples
For the MLEs which are sample means, the Bernoulli and Poisson MLEs from lecture 5, theorem 17.1 does not give any new results: the central limit theorem allows us to find the asymptotic normality of these estimators directly and from exercises 17.1 and 17.2 we know that the two variances are the same. For the normal mean with known variance, the MLE is exactly normally distributed for every , with variance . Thus, for this case the theorem gives an exact result. The first new case is an MLE which is not a sample mean.
Let be i.i.d. with exponential rate . Then the MLE for the parameter value is from example 5.2 and by exercise 16.1 the regularity conditions and the uniqueness of the root hold. From example 11.8 we know that the Fisher information is given by . Thus, by theorem 17.1,
with standard error . We can compare the result to exact results as follows: In lecture 5 we found , which is a bias of . In exercise 5.3 we found the variance . If we multiply this variance by , it converges to . If we multiply the bias by , it converges to zero, as predicted by the theorem: for large , the bias is of a smaller order than the spread of the distribution, and thus disappears in the limit.
The theorem can also be applied to functions of the parameter. The mean of the exponential distribution is and by theorem 5.6 the MLE for this parameter is . Using the delta method, theorem A.18 with , we get
This is the central limit theorem for , since the exponential distribution has variance . The standard error of , , is the standard error we would obtain by computing the standard error from first principles.
The example illustrates a general point: for any smooth function of the parameter, theorem 5.6 gives the MLE , and the delta method gives the asymptotic distribution and standard error , without any further maximisation.
17.5 Quality of the Approximation
Theorem 17.1 is a limit statement and does not give information about the required sample size for the normal approximation to be applicable. The required value of for the normal approximation to be usable depends on the model and there are three situations where care is needed. First, for small samples from a model with a skewed distribution: for the exponential rate, has a right-skewed distribution and a bias of which is not captured in the limit but which is visible for around . Secondly, when a parameter value is close to the boundary of the parameter space: for the Bernoulli distribution with close to zero or one, the standard error is zero whenever all observations have the same value, and thus the interval has zero width in this case. Finally, if the regularity conditions are violated: for the uniform distribution on , exercise 17.3 shows that the error of the MLE is of order instead of , and that the limiting distribution is exponential instead of normal, as expected from the super-efficiency of the parameter estimates as shown in lecture 13. The first case is investigated by simulation in lecture 12, by checking the coverage of the interval for moderate .
Finally, the theorem can be extended to a vector parameter : Under appropriate regularity conditions, the MLE is asymptotically normal with mean and covariance matrix , the inverse of the Fisher information matrix from lecture 14. Without proof, the argument is similar to the one above, but uses Taylor expansion for several variables. For the normal distribution with unknown parameters, where the information matrix is diagonal, this gives the asymptotic normality of and of the variance estimator separately.
-
•
Under regularity conditions and assuming the uniqueness of the likelihood root, we have : the MLE is asymptotically unbiased, the error is of order , and the asymptotic variance equals the Cramer–Rao bound, i.e. the MLE is asymptotically efficient.
-
•
The proof uses Taylor expansion of the score around , central limit theorem for the score, law of large numbers for the second derivative, consistency and (R4) for the remainder, and Slutsky’s theorem to combine these results.
-
•
Replacing by in the expression for the variance does not change the limit and thus we get the standard error and the approximate confidence interval for the parameter value.
-
•
Using the delta method, we can apply the result to any smooth function of the parameter. The standard error for such an application is .
-
•
The approximation can be poor for small in models with skewed distributions and at the boundary of the parameter space, and it breaks down completely when the regularity conditions are violated, e.g. for the uniform distribution.
Let be i.i.d. Bernoulli with success probability , where the MLE and Fisher information are given in example 11.5.
-
1.
State the result of theorem 17.1 for this model and verify that it coincides with the central limit theorem for .
-
2.
The log-odds are given by . Find the MLE for and use the delta method to show that .
-
3.
Determine the standard error of and an approximate interval for , in terms of and .
Let be i.i.d. Poisson with parameter , where MLE and Fisher information are given in example 11.6.
-
1.
State the result of theorem 17.1 for this model and verify that the result coincides with the central limit theorem.
-
2.
In exercise 5.7 we have found the MLE for the probability and we have seen that this MLE is biased. Using the delta method, determine the asymptotic distribution of . Why does the bias, found in the previous step, not contradict the result from this question?
Let be i.i.d. uniformly distributed on the interval . The MLE for this model is .
-
1.
For every , show that as , i.e. that converges in distribution to the exponential distribution with mean , i.e. with rate .
- 2.