Lecture 16 Consistency of the MLE
In lecture 5 we have seen that the maximum likelihood estimator can be biased, that it may not be unique and that it may not even exist. The following two lectures present the results which, despite these problems, allow us to justify the use of the MLE: for large sample size the maximum likelihood estimator converges to the true parameter value, and does so with minimal spread. In the current lecture we will prove the first of these two results, the consistency of the MLE, from a set of assumptions which we will explicitly state. The proof of the asymptotic normality in lecture 17 builds on these same assumptions and on the consistency result. The numerical aspects, concerning estimates and standard errors, will be discussed in lecture 18.
16.1 The Regularity Conditions
Throughout this and the following lectures, will be i.i.d. with density or probability weights , where the parameter is a one-dimensional parameter and is an open interval. The true parameter value is denoted by and and are probabilities and expectations w.r.t. this value. As in lecture 4, the log-likelihood is given by
and for random data the are i.i.d. random variables. The proofs in these two lectures will make use of the additive structure of the log-likelihood, allowing for applications of the law of large numbers and the central limit theorem for the likelihood.
The large-sample results are true for most models of interest, but there are exceptions; the uniform distribution from example 5.5 is a case in which the results do not hold. We start our discussion by stating the assumptions required for the theorems to hold.
The model satisfies the regularity conditions, if the following four statements hold.
-
(R1)
The parameter is identifiable in the sense of definition 4.10, i.e. different values of lead to different distributions.
-
(R2)
The support does not depend on .
-
(R3)
For every the function is three times continuously differentiable on , the order of differentiation with respect to and integration (or summation) with respect to can be interchanged twice, and the Fisher information of a single observation is a continuous function of with for all .
-
(R4)
For every there are a and a function such that and
The four conditions play different roles. Condition (R1) is needed for the question to make sense: if two parameter values give the same distribution, no estimator can converge to both. Condition (R2) excludes models like the uniform distribution, where the parameter moves the edge of the support and the likelihood jumps to zero; the Cramer–Rao bound fails for such models, as we have seen in lecture 13, and the arguments of this lecture fail for them too. Condition (R3) is the smoothness needed to differentiate the log-likelihood, and under (R2) and (R3) it makes the two identities of lecture 11 available (lemma 11.1 and theorem 11.3),
Condition (R4) is not needed in this lecture at all: it controls the remainder term in the Taylor expansion in the proof of asymptotic normality in lecture 17, and it is listed here so that both lectures work from the same definition. Although the condition looks technical, it is usually easy to verify, because the third derivative of the log-likelihood is an explicit function of and ; exercise 16.1 does this for the exponential distribution.
The usual families of distributions from tables 1.1 and 1.2, with the exception of the uniform distribution, satisfy the regularity conditions (with known for the binomial distribution). The same is true for the location model from section 5.3, despite the fact that the likelihood equation may have more than one solution. We will return to this problem in section 16.4.
16.2 The Likelihood Ratio Detects the Truth
If is the true value, it seems natural that for a large sample the likelihood should be larger than for any wrong value . This fact is the engine of the consistency proof, and thus we prove it first.
Suppose (R1) and (R2). Let and assume that has finite expectation under . Then
Using logarithms and dividing by , we can rewrite the event as
where, by (R2), the ratios are well defined and positive for all observations. The left-hand side is the sample mean of the i.i.d. random variables and by the law of large numbers, theorem A.14, converges in probability to . Therefore, we only need to show that this expectation is strictly negative, so that the sample mean is negative with probability approaching one.
Since the logarithm is strictly concave, the function is convex and we can use Jensen’s inequality, lemma A.7, to get
and the inequality is strict unless the ratio is constant with probability one. If the ratio was constant, we would have on the common support and, since both functions have integral one, this would imply , i.e. the two distributions would be the same which (R1) excludes for . Thus the inequality is strict. Writing the expectation on the right-hand side as an integral over the common support, we can cancel the density and get
with a sum instead of an integral for a discrete model. Thus we find and the proof is complete.
The assumption that is finite excludes nothing of interest: the only alternative is , and in that case the conclusion holds all the more easily. Replacing by for a constant gives i.i.d. variables with finite expectation, to which the law of large numbers applies, and that expectation is negative once is large enough; since the sample mean of the is at most that of the truncated variables, it is again negative with probability approaching one. ∎
The quantity from the proof is the Kullback–Leibler divergence of the wrong model from the true model, and the theorem states that on average each observation increases the log-likelihood of the truth by this fixed positive amount, compared to the log-likelihood of any of the competitors. It is important to note that the theorem only compares fixed values and does not state that will eventually be bigger than for all wrong , or even that it is a statement about the maximiser. In the next section we will see a proof that allows us to conclude that the statement holds for the maximiser.
16.3 Consistency
From definition 2.8 we know that an estimator is consistent for , if for all possible true values. In lecture 2 we have already seen how to prove consistency using the mean squared error. For the maximum likelihood estimator this approach is not feasible in general, because there is no general formula for the mean squared error. Instead, we will use the likelihood equation from lecture 5 to show that with probability tending to one there is a root of the equation which is close to and, if the root is unique, this root coincides with the MLE.
Suppose that the regularity conditions (R1), (R2) and (R3) hold, that has finite mean under for all , and that for every sample size and every sample with coordinates in the support of the model, which by (R2) is the same for every , the likelihood equation has a unique root which coincides with the maximum likelihood estimator . Then is a consistent estimator for .
Let the true value be and let be small enough that the interval is contained in the open set . Let and and consider the event
Since the event is the intersection of two events of the form considered in theorem 16.2, both of which have probability tending to one, we have as .
We now show that on the event the MLE is within of . Assume that the sample lies in and, to be specific, that ; the other case is symmetric. By (R3) the likelihood is a continuous function of and since the intermediate value theorem guarantees a point with . The function is differentiable and has the same value at the two ends of the interval , and thus by Rolle’s theorem there is a point with . Since on this interval, the point is also a root of the likelihood equation . By assumption, this root is unique and equals the MLE, and thus . By construction is in and we have shown that
Taking probabilities we find , i.e. as . Since was arbitrary and was an arbitrary point of , this is the statement for every true value, and the proof is complete. ∎
The structure of this proof will be re-used in lecture 17. The probabilistic input comes from theorem 16.2 which gives an event of probability tending to one; on this event the argument is then a standard application of calculus, and the uniqueness of the MLE allows us to pass from “some root of the likelihood equation” to “the MLE”. In the standard examples, uniqueness is typically proved by showing that the log-likelihood is strictly concave, as for the exponential distribution in example 5.2; exercise 16.1 completes the proof of this model.
The assumption that the likelihood equation has a unique root for every sample is slightly stronger than required. For example, for the Bernoulli distribution in example 5.3, the likelihood equation has no root in the case where all observations are equal, but this only happens with probability which tends to zero. The proof is unchanged if we assume that the condition holds on an event with probability tending to one, since we can intersect with such an event.
16.4 When the Assumptions Fail
Each hypothesis of theorem 16.3 can fail, and the failures are the cases where more care is needed in practice.
If the parameter is not identifiable, i.e. if (R1) fails, the question is ill posed: in example 4.11 the same distribution of the data could be generated by two different parameter values and thus no function of the data could converge to one of these parameter values rather than the other. In this case we would need to reparametrise the model, as shown in the example.
If the support of the model depends on the parameter, (R2), the likelihood is not a smooth function of and the MLE is typically not a root of the likelihood equation. For example, for the uniform distribution on , the MLE is the sample maximum, as shown in example 5.5, and the theorem does not apply in this case. Nevertheless, the estimator is consistent: the mean squared error of the estimator goes to zero as the sample size increases, as shown in exercise 2.2, and thus theorem 2.10 applies; exercise 16.3 gives a direct proof of the same fact. For the same model, in lecture 13, we have seen that the Cramer–Rao bound can be violated if (R2) fails and in lecture 17 we will see that the limiting distribution of the MLE is not normal in this case.
If the likelihood equation has several roots, the regularity conditions may all hold and the theorem still does not apply, because we cannot identify the root found by Rolle’s theorem with the MLE; the location model of section 5.3 is of this type. What survives is this: under (R1) to (R3) there is always a sequence of roots of the likelihood equation which is consistent, since the root found above lies within of . What is lost is the guarantee that the root we find, or the global maximiser, is this consistent one; we do not prove the general statement in this module. The practical consequence of this is that numerical maximisation of the likelihood should be started with a simple consistent estimator, such as the sample median for a location parameter, so that the iteration converges to the nearby consistent root. This is a lesson we will take up in lecture 18.
The following questions, about the rate of convergence of the MLE to and the distribution of the error for large , are answered in the asymptotic normality theorem of lecture 17. The proof of this theorem uses the consistency shown here, and condition (R4).
-
•
The regularity conditions (R1) to (R4) from definition 16.1 state that the parameters must be identifiable, the support of the model must not depend on , the log-likelihood must be three times differentiable with finite and positive Fisher information, and a bound must be placed on the third derivative of the log-likelihood.
-
•
For fixed parameter values, the likelihood of the true value eventually exceeds the likelihood of any wrong value, with probability approaching one. The proof uses the law of large numbers and Jensen’s inequality.
-
•
If the likelihood equation has only one root, and this root is the MLE, then the MLE is consistent. The proof finds a root of the likelihood equation near , using Rolle’s theorem and the fact that this root coincides with the MLE.
-
•
The theorem does not apply if the support of the model depends on the parameter (since then the MLE is typically not a root of the likelihood equation, as for the uniform distribution, even though it may still be consistent), or if the likelihood equation has multiple roots (since then only the existence of a consistent root can be proved).
Let be i.i.d. with exponential rate , i.e. for all .
-
1.
Show that the regularity conditions (R1) to (R4) hold for this model. Hint: For (R3) you can use the value from example 11.8. For (R4) you should take derivatives of with respect to and then find a bound .
-
2.
Show that the likelihood equation has only one root for each sample and that this root is the MLE. Use theorem 16.3 to conclude that is consistent for .
- 3.
Let be the only observation of a Bernoulli random variable with success probability , and let be another parameter value.
-
1.
Compute the expectation from the proof of theorem 16.2, as an explicit function of and .
-
2.
Using the inequality for , with equality only for , show that this expectation is negative for all .
Let be i.i.d. uniformly distributed on the interval and let be the MLE from example 5.5.
-
1.
Show that for we have and thus is consistent for .
-
2.
Which of the assumptions of theorem 16.3 are violated for this model? Why cannot the theorem be applied in this case, despite the fact that the conclusion is still true?