Lecture 16 Consistency of the MLE

In lecture 5 we have seen that the maximum likelihood estimator can be biased, that it may not be unique and that it may not even exist. The following two lectures present the results which, despite these problems, allow us to justify the use of the MLE: for large sample size the maximum likelihood estimator converges to the true parameter value, and does so with minimal spread. In the current lecture we will prove the first of these two results, the consistency of the MLE, from a set of assumptions which we will explicitly state. The proof of the asymptotic normality in lecture 17 builds on these same assumptions and on the consistency result. The numerical aspects, concerning estimates and standard errors, will be discussed in lecture 18.

16.1 The Regularity Conditions

Throughout this and the following lectures, X1,…,XnX_{1},\dots,X_{n} will be i.i.d. with density or probability weights f⁢(x;θ)f(x;\theta), where the parameter θ\theta is a one-dimensional parameter and Θ⊆ℝ\Theta\subseteq\mathbb{R} is an open interval. The true parameter value is denoted by θ0\theta_{0} and ℙθ0\mathbb{P}_{\theta_{0}} and 𝔼θ0\mathbb{E}_{\theta_{0}} are probabilities and expectations w.r.t. this value. As in lecture 4, the log-likelihood is given by

ℓ⁢(θ)=∑i=1nℓi⁢(θ),where ⁢ℓi⁢(θ)=log⁡f⁢(Xi;θ),\ell(\theta)=\sum_{i=1}^{n}\ell_{i}(\theta),\qquad\text{where }\ell_{i}(\theta% )=\log f(X_{i};\theta),

and for random data the ℓi⁢(θ)\ell_{i}(\theta) are i.i.d. random variables. The proofs in these two lectures will make use of the additive structure of the log-likelihood, allowing for applications of the law of large numbers and the central limit theorem for the likelihood.

The large-sample results are true for most models of interest, but there are exceptions; the uniform distribution from example 5.5 is a case in which the results do not hold. We start our discussion by stating the assumptions required for the theorems to hold.

Definition 16.1.

The model (f⁢(x;θ)|θ∈Θ)\bigl{(}f(x;\theta)\!\mathrel{\big{|}}\!\theta\in\Theta\bigr{)} satisfies the regularity conditions, if the following four statements hold.

  1. (R1)
    ​

    The parameter is identifiable in the sense of definition 4.10, i.e. different values of θ\theta lead to different distributions.

  2. (R2)
    ​

    The support {x∣f⁢(x;θ)>0}\{x\mid f(x;\theta)>0\} does not depend on θ\theta.

  3. (R3)
    ​

    For every xx the function θ↦f⁢(x;θ)\theta\mapsto f(x;\theta) is three times continuously differentiable on Θ\Theta, the order of differentiation with respect to θ\theta and integration (or summation) with respect to xx can be interchanged twice, and the Fisher information ℐ︀⁢(θ)\mathcal{I}(\theta) of a single observation is a continuous function of θ\theta with 0<ℐ︀⁢(θ)<∞0<\mathcal{I}(\theta)<\infty for all θ∈Θ\theta\in\Theta.

  4. (R4)
    ​

    For every θ0∈Θ\theta_{0}\in\Theta there are a δ>0\delta>0 and a function M⁢(x)M(x) such that 𝔼θ0⁢(M⁢(X1))<∞\mathbb{E}_{\theta_{0}}\bigl{(}M(X_{1})\bigr{)}<\infty and

    |∂3∂θ3⁢log⁡f⁢(x;θ)|≤M⁢(x)for all x and all θ with |θ−θ0|<δ.\Bigl{|}\frac{\partial^{3}}{\partial\theta^{3}}\log f(x;\theta)\Bigr{|}\leq M(% x)\qquad\text{for all $x$ and all $\theta$ with $|\theta-\theta_{0}|<\delta$.}

The four conditions play different roles. Condition (R1) is needed for the question to make sense: if two parameter values give the same distribution, no estimator can converge to both. Condition (R2) excludes models like the uniform distribution, where the parameter moves the edge of the support and the likelihood jumps to zero; the Cramer–Rao bound fails for such models, as we have seen in lecture 13, and the arguments of this lecture fail for them too. Condition (R3) is the smoothness needed to differentiate the log-likelihood, and under (R2) and (R3) it makes the two identities of lecture 11 available (lemma 11.1 and theorem 11.3),

equation (16.1) (16.1)
𝔼θ⁢(ℓi′⁢(θ))=0andℐ︀⁢(θ)=Varθ(ℓi′⁢(θ))=−𝔼θ⁢(ℓi′′⁢(θ)).\mathbb{E}_{\theta}\bigl{(}\ell_{i}^{\prime}(\theta)\bigr{)}=0\qquad\text{and}% \qquad\mathcal{I}(\theta)=\mathop{\mathrm{Var}}\nolimits_{\theta}\bigl{(}\ell_% {i}^{\prime}(\theta)\bigr{)}=-\mathbb{E}_{\theta}\bigl{(}\ell_{i}^{\prime% \prime}(\theta)\bigr{)}.

Condition (R4) is not needed in this lecture at all: it controls the remainder term in the Taylor expansion in the proof of asymptotic normality in lecture 17, and it is listed here so that both lectures work from the same definition. Although the condition looks technical, it is usually easy to verify, because the third derivative of the log-likelihood is an explicit function of xx and θ\theta; exercise 16.1 does this for the exponential distribution.

The usual families of distributions from tables 1.1 and 1.2, with the exception of the uniform distribution, satisfy the regularity conditions (with mm known for the binomial distribution). The same is true for the tt location model from section 5.3, despite the fact that the likelihood equation may have more than one solution. We will return to this problem in section 16.4.

16.2 The Likelihood Ratio Detects the Truth

If θ0\theta_{0} is the true value, it seems natural that for a large sample the likelihood L⁢(θ0)L(\theta_{0}) should be larger than L⁢(θ)L(\theta) for any wrong value θ\theta. This fact is the engine of the consistency proof, and thus we prove it first.

Theorem 16.2.

Suppose (R1) and (R2). Let θ≠θ0\theta\neq\theta_{0} and assume that Y1=log⁡(f⁢(X1;θ)/f⁢(X1;θ0))Y_{1}=\log\bigl{(}f(X_{1};\theta)/f(X_{1};\theta_{0})\bigr{)} has finite expectation under ℙθ0\mathbb{P}_{\theta_{0}}. Then

limn→∞ℙθ0⁢(L⁢(θ0)>L⁢(θ))=1.\lim_{n\to\infty}\mathbb{P}_{\theta_{0}}\bigl{(}L(\theta_{0})>L(\theta)\bigr{)% }=1.
Proof.

Using logarithms and dividing by nn, we can rewrite the event L⁢(θ0)>L⁢(θ)L(\theta_{0})>L(\theta) as

1n⁢∑i=1nlog⁡f⁢(Xi;θ)f⁢(Xi;θ0)<0,\frac{1}{n}\sum_{i=1}^{n}\log\frac{f(X_{i};\theta)}{f(X_{i};\theta_{0})}<0,

where, by (R2), the ratios are well defined and positive for all observations. The left-hand side is the sample mean of the i.i.d. random variables Yi=log⁡(f⁢(Xi;θ)/f⁢(Xi;θ0))Y_{i}=\log\bigl{(}f(X_{i};\theta)/f(X_{i};\theta_{0})\bigr{)} and by the law of large numbers, theorem A.14, converges in probability to 𝔼θ0⁢(Y1)\mathbb{E}_{\theta_{0}}(Y_{1}). Therefore, we only need to show that this expectation is strictly negative, so that the sample mean is negative with probability approaching one.

Since the logarithm is strictly concave, the function −log-\log is convex and we can use Jensen’s inequality, lemma A.7, to get

𝔼θ0⁢(Y1)=𝔼θ0⁢(log⁡f⁢(X1;θ)f⁢(X1;θ0))≤log⁡𝔼θ0⁢(f⁢(X1;θ)f⁢(X1;θ0)),\mathbb{E}_{\theta_{0}}(Y_{1})=\mathbb{E}_{\theta_{0}}\Bigl{(}\log\frac{f(X_{1% };\theta)}{f(X_{1};\theta_{0})}\Bigr{)}\leq\log\mathbb{E}_{\theta_{0}}\Bigl{(}% \frac{f(X_{1};\theta)}{f(X_{1};\theta_{0})}\Bigr{)},

and the inequality is strict unless the ratio f⁢(X1;θ)/f⁢(X1;θ0)f(X_{1};\theta)/f(X_{1};\theta_{0}) is constant with probability one. If the ratio was constant, we would have f⁢(x;θ)=c⁢f⁢(x;θ0)f(x;\theta)=c\,f(x;\theta_{0}) on the common support and, since both functions have integral one, this would imply c=1c=1, i.e. the two distributions would be the same which (R1) excludes for θ≠θ0\theta\neq\theta_{0}. Thus the inequality is strict. Writing the expectation on the right-hand side as an integral over the common support, we can cancel the density f⁢(x;θ0)f(x;\theta_{0}) and get

𝔼θ0⁢(f⁢(X1;θ)f⁢(X1;θ0))=∫f⁢(x;θ)f⁢(x;θ0)⁢f⁢(x;θ0)⁢dx=∫f⁢(x;θ)⁢dx=1,\mathbb{E}_{\theta_{0}}\Bigl{(}\frac{f(X_{1};\theta)}{f(X_{1};\theta_{0})}% \Bigr{)}=\int\frac{f(x;\theta)}{f(x;\theta_{0})}\,f(x;\theta_{0})\,\mathrm{d}x% =\int f(x;\theta)\,\mathrm{d}x=1,

with a sum instead of an integral for a discrete model. Thus we find 𝔼θ0⁢(Y1)<log⁡1=0\mathbb{E}_{\theta_{0}}(Y_{1})<\log 1=0 and the proof is complete.

The assumption that 𝔼θ0⁢(Y1)\mathbb{E}_{\theta_{0}}(Y_{1}) is finite excludes nothing of interest: the only alternative is 𝔼θ0⁢(Y1)=−∞\mathbb{E}_{\theta_{0}}(Y_{1})=-\infty, and in that case the conclusion holds all the more easily. Replacing YiY_{i} by max⁡(Yi,−M)\max(Y_{i},-M) for a constant M>0M>0 gives i.i.d. variables with finite expectation, to which the law of large numbers applies, and that expectation is negative once MM is large enough; since the sample mean of the YiY_{i} is at most that of the truncated variables, it is again negative with probability approaching one. ∎

The quantity −𝔼θ0⁢(Y1)-\mathbb{E}_{\theta_{0}}(Y_{1}) from the proof is the Kullback–Leibler divergence of the wrong model from the true model, and the theorem states that on average each observation increases the log-likelihood of the truth by this fixed positive amount, compared to the log-likelihood of any of the competitors. It is important to note that the theorem only compares fixed values and does not state that L⁢(θ0)L(\theta_{0}) will eventually be bigger than L⁢(θ)L(\theta) for all wrong θ\theta, or even that it is a statement about the maximiser. In the next section we will see a proof that allows us to conclude that the statement holds for the maximiser.

16.3 Consistency

From definition 2.8 we know that an estimator θ^n\hat{\theta}_{n} is consistent for θ\theta, if θ^n→pθ\hat{\theta}_{n}\xrightarrow{\ \mathrm{p}\ }\theta for all possible true values. In lecture 2 we have already seen how to prove consistency using the mean squared error. For the maximum likelihood estimator this approach is not feasible in general, because there is no general formula for the mean squared error. Instead, we will use the likelihood equation from lecture 5 to show that with probability tending to one there is a root of the equation which is close to θ0\theta_{0} and, if the root is unique, this root coincides with the MLE.

Theorem 16.3.

Suppose that the regularity conditions (R1), (R2) and (R3) hold, that log⁡(f⁢(X1;θ)/f⁢(X1;θ′))\log\bigl{(}f(X_{1};\theta)/f(X_{1};\theta^{\prime})\bigr{)} has finite mean under ℙθ′\mathbb{P}_{\theta^{\prime}} for all θ,θ′∈Θ\theta,\theta^{\prime}\in\Theta, and that for every sample size n≥1n\geq 1 and every sample with coordinates in the support of the model, which by (R2) is the same for every θ\theta, the likelihood equation ℓ′⁢(θ)=0\ell^{\prime}(\theta)=0 has a unique root which coincides with the maximum likelihood estimator θ^n\hat{\theta}_{n}. Then θ^n\hat{\theta}_{n} is a consistent estimator for θ\theta.

Proof.

Let the true value be θ0\theta_{0} and let ε>0\varepsilon>0 be small enough that the interval (θ0−ε,θ0+ε)(\theta_{0}-\varepsilon,\theta_{0}+\varepsilon) is contained in the open set Θ\Theta. Let θ1=θ0−ε\theta_{1}=\theta_{0}-\varepsilon and θ2=θ0+ε\theta_{2}=\theta_{0}+\varepsilon and consider the event

Sε={L⁢(θ1)<L⁢(θ0)⁢ and ⁢L⁢(θ2)<L⁢(θ0)}.S_{\varepsilon}=\bigl{\{}L(\theta_{1})<L(\theta_{0})\text{ and }L(\theta_{2})<% L(\theta_{0})\bigr{\}}.

Since the event SεS_{\varepsilon} is the intersection of two events of the form considered in theorem 16.2, both of which have probability tending to one, we have ℙθ0⁢(Sε)→1\mathbb{P}_{\theta_{0}}(S_{\varepsilon})\to 1 as n→∞n\to\infty.

We now show that on the event SεS_{\varepsilon} the MLE is within ε\varepsilon of θ0\theta_{0}. Assume that the sample lies in SεS_{\varepsilon} and, to be specific, that L⁢(θ1)≤L⁢(θ2)L(\theta_{1})\leq L(\theta_{2}); the other case is symmetric. By (R3) the likelihood is a continuous function of θ\theta and since L⁢(θ1)≤L⁢(θ2)<L⁢(θ0)L(\theta_{1})\leq L(\theta_{2})<L(\theta_{0}) the intermediate value theorem guarantees a point θ~1∈[θ1,θ0)\tilde{\theta}_{1}\in[\theta_{1},\theta_{0}) with L⁢(θ~1)=L⁢(θ2)L(\tilde{\theta}_{1})=L(\theta_{2}). The function LL is differentiable and has the same value at the two ends of the interval [θ~1,θ2][\tilde{\theta}_{1},\theta_{2}], and thus by Rolle’s theorem there is a point θ∗∈(θ~1,θ2)\theta^{*}\in(\tilde{\theta}_{1},\theta_{2}) with L′⁢(θ∗)=0L^{\prime}(\theta^{*})=0. Since L>0L>0 on this interval, the point θ∗\theta^{*} is also a root of the likelihood equation ℓ′⁢(θ)=0\ell^{\prime}(\theta)=0. By assumption, this root is unique and equals the MLE, and thus θ^n=θ∗\hat{\theta}_{n}=\theta^{*}. By construction θ∗\theta^{*} is in (θ1,θ2)(\theta_{1},\theta_{2}) and we have shown that

Sε⊆{|θ^n−θ0|<ε}.S_{\varepsilon}\subseteq\bigl{\{}|\hat{\theta}_{n}-\theta_{0}|<\varepsilon% \bigr{\}}.

Taking probabilities we find ℙθ0⁢(|θ^n−θ0|<ε)≥ℙθ0⁢(Sε)→1\mathbb{P}_{\theta_{0}}\bigl{(}|\hat{\theta}_{n}-\theta_{0}|<\varepsilon\bigr{% )}\geq\mathbb{P}_{\theta_{0}}(S_{\varepsilon})\to 1, i.e. ℙθ0⁢(|θ^n−θ0|≥ε)→0\mathbb{P}_{\theta_{0}}\bigl{(}|\hat{\theta}_{n}-\theta_{0}|\geq\varepsilon% \bigr{)}\to 0 as n→∞n\to\infty. Since ε>0\varepsilon>0 was arbitrary and θ0\theta_{0} was an arbitrary point of Θ\Theta, this is the statement θ^n→pθ\hat{\theta}_{n}\xrightarrow{\ \mathrm{p}\ }\theta for every true value, and the proof is complete. ∎

The structure of this proof will be re-used in lecture 17. The probabilistic input comes from theorem 16.2 which gives an event of probability tending to one; on this event the argument is then a standard application of calculus, and the uniqueness of the MLE allows us to pass from “some root of the likelihood equation” to “the MLE”. In the standard examples, uniqueness is typically proved by showing that the log-likelihood is strictly concave, as for the exponential distribution in example 5.2; exercise 16.1 completes the proof of this model.

Remark.

The assumption that the likelihood equation has a unique root for every sample is slightly stronger than required. For example, for the Bernoulli distribution in example 5.3, the likelihood equation has no root in the case where all observations are equal, but this only happens with probability θ0n+(1−θ0)n\theta_{0}^{n}+(1-\theta_{0})^{n} which tends to zero. The proof is unchanged if we assume that the condition holds on an event with probability tending to one, since we can intersect SεS_{\varepsilon} with such an event.

16.4 When the Assumptions Fail

Each hypothesis of theorem 16.3 can fail, and the failures are the cases where more care is needed in practice.

If the parameter is not identifiable, i.e. if (R1) fails, the question is ill posed: in example 4.11 the same distribution of the data could be generated by two different parameter values and thus no function of the data could converge to one of these parameter values rather than the other. In this case we would need to reparametrise the model, as shown in the example.

If the support of the model depends on the parameter, (R2), the likelihood is not a smooth function of θ\theta and the MLE is typically not a root of the likelihood equation. For example, for the uniform distribution on (0,θ)(0,\theta), the MLE is the sample maximum, as shown in example 5.5, and the theorem does not apply in this case. Nevertheless, the estimator is consistent: the mean squared error of the estimator goes to zero as the sample size increases, as shown in exercise 2.2, and thus theorem 2.10 applies; exercise 16.3 gives a direct proof of the same fact. For the same model, in lecture 13, we have seen that the Cramer–Rao bound can be violated if (R2) fails and in lecture 17 we will see that the limiting distribution of the MLE is not normal in this case.

If the likelihood equation has several roots, the regularity conditions may all hold and the theorem still does not apply, because we cannot identify the root found by Rolle’s theorem with the MLE; the tt location model of section 5.3 is of this type. What survives is this: under (R1) to (R3) there is always a sequence of roots of the likelihood equation which is consistent, since the root θ∗\theta^{*} found above lies within ε\varepsilon of θ0\theta_{0}. What is lost is the guarantee that the root we find, or the global maximiser, is this consistent one; we do not prove the general statement in this module. The practical consequence of this is that numerical maximisation of the likelihood should be started with a simple consistent estimator, such as the sample median for a location parameter, so that the iteration converges to the nearby consistent root. This is a lesson we will take up in lecture 18.

The following questions, about the rate of convergence of the MLE to θ0\theta_{0} and the distribution of the error for large nn, are answered in the asymptotic normality theorem of lecture 17. The proof of this theorem uses the consistency shown here, and condition (R4).

Summary.
  • •
    ​

    The regularity conditions (R1) to (R4) from definition 16.1 state that the parameters must be identifiable, the support of the model must not depend on θ\theta, the log-likelihood must be three times differentiable with finite and positive Fisher information, and a bound must be placed on the third derivative of the log-likelihood.

  • •
    ​

    For fixed parameter values, the likelihood of the true value eventually exceeds the likelihood of any wrong value, with probability approaching one. The proof uses the law of large numbers and Jensen’s inequality.

  • •
    ​

    If the likelihood equation has only one root, and this root is the MLE, then the MLE is consistent. The proof finds a root of the likelihood equation near θ0\theta_{0}, using Rolle’s theorem and the fact that this root coincides with the MLE.

  • •
    ​

    The theorem does not apply if the support of the model depends on the parameter (since then the MLE is typically not a root of the likelihood equation, as for the uniform distribution, even though it may still be consistent), or if the likelihood equation has multiple roots (since then only the existence of a consistent root can be proved).

Exercise 16.1.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with exponential rate θ>0\theta>0, i.e. f⁢(x;θ)=θ⁢e−θ⁢xf(x;\theta)=\theta e^{-\theta x} for all x≥0x\geq 0.

  1. 1.
    ​

    Show that the regularity conditions (R1) to (R4) hold for this model. Hint: For (R3) you can use the value ℐ︀⁢(θ)=1/θ2\mathcal{I}(\theta)=1/\theta^{2} from example 11.8. For (R4) you should take derivatives of log⁡f⁢(x;θ)\log f(x;\theta) with respect to θ\theta and then find a bound M⁢(x)M(x).

  2. 2.
    ​

    Show that the likelihood equation has only one root for each sample and that this root is the MLE. Use theorem 16.3 to conclude that θ^n=1/X¯\hat{\theta}_{n}=1/\bar{X} is consistent for θ\theta.

  3. 3.
    ​

    Find an alternative proof of the consistency of 1/X¯1/\bar{X}, using the law of large numbers (theorem A.14), the continuous mapping theorem (theorem A.17), and any other arguments you need.

Exercise 16.2.

Let X1X_{1} be the only observation of a Bernoulli random variable with success probability θ0∈(0,1)\theta_{0}\in(0,1), and let θ∈(0,1)\theta\in(0,1) be another parameter value.

  1. 1.
    ​

    Compute the expectation 𝔼θ0⁢(log⁡(f⁢(X1;θ)/f⁢(X1;θ0)))\mathbb{E}_{\theta_{0}}\bigl{(}\log\bigl{(}f(X_{1};\theta)/f(X_{1};\theta_{0})% \bigr{)}\bigr{)} from the proof of theorem 16.2, as an explicit function of θ\theta and θ0\theta_{0}.

  2. 2.
    ​

    Using the inequality log⁡y≤y−1\log y\leq y-1 for y>0y>0, with equality only for y=1y=1, show that this expectation is negative for all θ≠θ0\theta\neq\theta_{0}.

Exercise 16.3.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. uniformly distributed on the interval (0,θ)(0,\theta) and let θ^n=maxi⁡Xi\hat{\theta}_{n}=\max_{i}X_{i} be the MLE from example 5.5.

  1. 1.
    ​

    Show that for 0<ε<θ0<\varepsilon<\theta we have ℙθ⁢(|θ^n−θ|>ε)=(1−ε/θ)n\mathbb{P}_{\theta}\bigl{(}|\hat{\theta}_{n}-\theta|>\varepsilon\bigr{)}=(1-% \varepsilon/\theta)^{n} and thus θ^n\hat{\theta}_{n} is consistent for θ\theta.

  2. 2.
    ​

    Which of the assumptions of theorem 16.3 are violated for this model? Why cannot the theorem be applied in this case, despite the fact that the conclusion is still true?