Lecture 17 Asymptotic Normality and Efficiency of the MLE

In lecture 16 we have seen that under regularity conditions the maximum likelihood estimator converges to the true parameter value. In this lecture we will discuss the error θ^n−θ0\hat{\theta}_{n}-\theta_{0} in more detail: It transpires that for large sample size the error is approximately normally distributed with mean zero and variance given by the Cramer–Rao bound from lecture 13. Thus, the MLE is unbiased in the limit, and more importantly, is as accurate as any unbiased estimator can be. The proof of this result is a challenging one, but it is based on tools we have already learned in this module. The argument is based on Taylor expansion of the score function, and the resulting theorem allows us to compute standard errors and to derive approximate confidence statements. We will see the resulting numbers in lecture 18 and will make the statements precise in lecture 25.

17.1 The Theorem

Using the notation from lecture 16, we write ℐ︀⁢(θ)\mathcal{I}(\theta) for the Fisher information of a single observation and ℐ︀n⁢(θ)=n⁢ℐ︀⁢(θ)\mathcal{I}_{n}(\theta)=n\mathcal{I}(\theta) for the Fisher information of the sample, as given in definition 11.2 and proposition 11.4. The result is formulated in terms of the scaled error n⁢(θ^n−θ0)\sqrt{n}\,(\hat{\theta}_{n}-\theta_{0}), as this is the only error term which has a non-degenerate limit: by consistency the error goes to zero and the factor n\sqrt{n} is required to balance this effect.

Theorem 17.1.

Suppose that the regularity conditions (R1) to (R4) from definition 16.1 are satisfied, that log⁡(f⁢(X1;θ)/f⁢(X1;θ′))\log\bigl{(}f(X_{1};\theta)/f(X_{1};\theta^{\prime})\bigr{)} has finite mean under ℙθ′\mathbb{P}_{\theta^{\prime}} for all θ,θ′∈Θ\theta,\theta^{\prime}\in\Theta, and that for every sample size n≥1n\geq 1 and every sample where all coordinates are in the support of the model, the likelihood equation ℓ′⁢(θ)=0\ell^{\prime}(\theta)=0 has a unique solution θ^n\hat{\theta}_{n} which coincides with the maximum likelihood estimator. Then

n⁢(θ^n−θ0)→dN⁢(0,1ℐ︀⁢(θ0))(n→∞).\sqrt{n}\,(\hat{\theta}_{n}-\theta_{0})\xrightarrow{\ \mathrm{d}\ }N\Bigl{(}0,% \frac{1}{\mathcal{I}(\theta_{0})}\Bigr{)}\qquad(n\to\infty).

Before we give the proof, we unpack the statement of the theorem. The theorem informally states that, as nn gets large,

θ^n≈N⁢(θ0,1n⁢ℐ︀⁢(θ0))=N⁢(θ0,1ℐ︀n⁢(θ0)).\hat{\theta}_{n}\approx N\Bigl{(}\theta_{0},\frac{1}{n\mathcal{I}(\theta_{0})}% \Bigr{)}=N\Bigl{(}\theta_{0},\frac{1}{\mathcal{I}_{n}(\theta_{0})}\Bigr{)}.

This statement contains three different aspects: (1) The centre of the approximating distribution equals the true value, i.e. the MLE is asymptotically unbiased (this does not mean that the MLE is unbiased for finite nn, only that it is asymptotically unbiased). (2) The spread of the distribution decreases as 1/n1/\sqrt{n}, as for the sample mean. (3) The variance of the approximating distribution equals 1/ℐ︀n⁢(θ0)1/\mathcal{I}_{n}(\theta_{0}). This is the Cramer–Rao lower bound from lecture 13: no unbiased estimator can have smaller variance for any nn and the MLE achieves this bound in the limit. The last property has a name.

Definition 17.2.

An estimator θ~n\tilde{\theta}_{n} for θ\theta, which does not have to be the maximum likelihood estimator, is asymptotically efficient, if n⁢(θ~n−θ)→dN⁢(0,1/ℐ︀⁢(θ))\sqrt{n}\,(\tilde{\theta}_{n}-\theta)\xrightarrow{\ \mathrm{d}\ }N\bigl{(}0,1/% \mathcal{I}(\theta)\bigr{)} for all θ∈Θ\theta\in\Theta.

In this language theorem 17.1 states that under the regularity conditions the MLE is asymptotically efficient. In lecture 13 we have seen that an efficient estimator, one which achieves the Cramer–Rao bound exactly, only exists for special models and parametrisations, namely the natural statistic of an exponential family. The asymptotic statement is much wider in applicability, as it holds for every regular model and every parametrisation, since by theorem 5.6 the MLE transforms together with the parameter.

17.2 Proof of the Theorem

We will use a Taylor expansion of the score function ℓ′\ell^{\prime} around the true value. Since the MLE is a root of the score, we can use Taylor expansion of ℓ′⁢(θ^n)\ell^{\prime}(\hat{\theta}_{n}) around the point θ0\theta_{0} to write the error term θ^n−θ0\hat{\theta}_{n}-\theta_{0} in terms of ℓ′⁢(θ0)\ell^{\prime}(\theta_{0}) and ℓ′′⁢(θ0)\ell^{\prime\prime}(\theta_{0}), which are both sums of i.i.d. terms and thus can be analysed using the limit theorems for sample means.

Proof of theorem 17.1.

Let the true value be θ0\theta_{0}. Then, by (R3), the function ℓ′\ell^{\prime} is two times continuously differentiable. Thus, by Taylor’s theorem with the Lagrange form of the remainder, for ℓ′\ell^{\prime} around the point θ0\theta_{0} we get

ℓ′⁢(θ^n)=ℓ′⁢(θ0)+ℓ′′⁢(θ0)⁢(θ^n−θ0)+12⁢ℓ′′′⁢(θn∗)⁢(θ^n−θ0)2,\ell^{\prime}(\hat{\theta}_{n})=\ell^{\prime}(\theta_{0})+\ell^{\prime\prime}(% \theta_{0})\,(\hat{\theta}_{n}-\theta_{0})+\tfrac{1}{2}\,\ell^{\prime\prime% \prime}(\theta_{n}^{*})\,(\hat{\theta}_{n}-\theta_{0})^{2},

where θn∗\theta_{n}^{*} is a point between θ0\theta_{0} and θ^n\hat{\theta}_{n}. Since θ^n\hat{\theta}_{n} solves the likelihood equation, the left-hand side is zero. Solving for θ^n−θ0\hat{\theta}_{n}-\theta_{0} and multiplying by n\sqrt{n} we can write the scaled error as a ratio:

equation (17.1) (17.1)
n⁢(θ^n−θ0)=AnBn−Cn,\sqrt{n}\,(\hat{\theta}_{n}-\theta_{0})=\frac{A_{n}}{B_{n}-C_{n}},

where

An=1n⁢ℓ′⁢(θ0),Bn=−1n⁢ℓ′′⁢(θ0),Cn=12⁢n⁢ℓ′′′⁢(θn∗)⁢(θ^n−θ0).A_{n}=\frac{1}{\sqrt{n}}\,\ell^{\prime}(\theta_{0}),\qquad B_{n}=-\frac{1}{n}% \,\ell^{\prime\prime}(\theta_{0}),\qquad C_{n}=\frac{1}{2n}\,\ell^{\prime% \prime\prime}(\theta_{n}^{*})\,(\hat{\theta}_{n}-\theta_{0}).

We now take the limit of each of the three terms separately.

The numerator AnA_{n} is a scaled sum of i.i.d. terms: by definition of the score we have ℓ′⁢(θ0)=∑iℓi′⁢(θ0)\ell^{\prime}(\theta_{0})=\sum_{i}\ell_{i}^{\prime}(\theta_{0}) and using the identities (16.1) we find that each term has mean zero and finite, positive variance ℐ︀⁢(θ0)\mathcal{I}(\theta_{0}). Thus, by the central limit theorem, theorem A.15, for the sample mean of the ℓi′⁢(θ0)\ell_{i}^{\prime}(\theta_{0}) we get

An=n⁢(1n⁢∑i=1nℓi′⁢(θ0)−0)→dN⁢(0,ℐ︀⁢(θ0)).A_{n}=\sqrt{n}\,\Bigl{(}\frac{1}{n}\sum_{i=1}^{n}\ell_{i}^{\prime}(\theta_{0})% -0\Bigr{)}\xrightarrow{\ \mathrm{d}\ }N\bigl{(}0,\mathcal{I}(\theta_{0})\bigr{% )}.

The first term in the denominator is also a sample mean: Bn=−1n⁢∑iℓi′′⁢(θ0)B_{n}=-\frac{1}{n}\sum_{i}\ell_{i}^{\prime\prime}(\theta_{0}) and by the law of large numbers, theorem A.14, and the second identity in (16.1) we find

Bn→p−𝔼θ0⁢(ℓ1′′⁢(θ0))=ℐ︀⁢(θ0).B_{n}\xrightarrow{\ \mathrm{p}\ }-\mathbb{E}_{\theta_{0}}\bigl{(}\ell_{1}^{% \prime\prime}(\theta_{0})\bigr{)}=\mathcal{I}(\theta_{0}).

The remainder term CnC_{n} is where the consistency from lecture 16 and condition (R4) come into play. Let δ\delta and MM be as in (R4). On the event {|θ^n−θ0|<δ}\{|\hat{\theta}_{n}-\theta_{0}|<\delta\}, the intermediate point θn∗\theta_{n}^{*} also lies within δ\delta of θ0\theta_{0}. Thus, writing ℓ′′′⁢(θn∗)=∑iℓi′′′⁢(θn∗)\ell^{\prime\prime\prime}(\theta_{n}^{*})=\sum_{i}\ell_{i}^{\prime\prime\prime% }(\theta_{n}^{*}) as a sum of the observations and using (R4) for each term, we find

|1n⁢ℓ′′′⁢(θn∗)|≤1n⁢∑i=1n|ℓi′′′⁢(θn∗)|≤1n⁢∑i=1nM⁢(Xi)\Bigl{|}\frac{1}{n}\,\ell^{\prime\prime\prime}(\theta_{n}^{*})\Bigr{|}\leq% \frac{1}{n}\sum_{i=1}^{n}\bigl{|}\ell_{i}^{\prime\prime\prime}(\theta_{n}^{*})% \bigr{|}\leq\frac{1}{n}\sum_{i=1}^{n}M(X_{i})

on this event. By theorem 16.3, the probability of the event goes to one, and by the law of large numbers the right-hand side converges in probability to the finite constant 𝔼θ0⁢(M⁢(X1))\mathbb{E}_{\theta_{0}}\bigl{(}M(X_{1})\bigr{)}. Thus, for all sufficiently large KK we have ℙθ0⁢(|ℓ′′′⁢(θn∗)|/n>K)→0\mathbb{P}_{\theta_{0}}\bigl{(}|\ell^{\prime\prime\prime}(\theta_{n}^{*})|/n>K% \bigr{)}\to 0. The other factor of CnC_{n} is θ^n−θ0\hat{\theta}_{n}-\theta_{0}, which converges in probability to zero by theorem 16.3. For any η>0\eta>0, the event |Cn|>η|C_{n}|>\eta is contained in the union of the events |ℓ′′′⁢(θn∗)|/n>K|\ell^{\prime\prime\prime}(\theta_{n}^{*})|/n>K and |θ^n−θ0|>2⁢η/K|\hat{\theta}_{n}-\theta_{0}|>2\eta/K, both of which have probability that goes to zero, and thus Cn→p0C_{n}\xrightarrow{\ \mathrm{p}\ }0.

It remains to combine these three limits, and the appropriate tool for this is Slutsky’s theorem, theorem A.16. Since Bn→pℐ︀⁢(θ0)B_{n}\xrightarrow{\ \mathrm{p}\ }\mathcal{I}(\theta_{0}) and Cn→p0C_{n}\xrightarrow{\ \mathrm{p}\ }0, the denominator satisfies Bn−Cn→pℐ︀⁢(θ0)B_{n}-C_{n}\xrightarrow{\ \mathrm{p}\ }\mathcal{I}(\theta_{0}), where the value is constant and non-zero. By the quotient part of Slutsky’s theorem, for the representation (17.1), we have

n⁢(θ^n−θ0)=AnBn−Cn→dN⁢(0,ℐ︀⁢(θ0))ℐ︀⁢(θ0)=N⁢(0,1ℐ︀⁢(θ0)),\sqrt{n}\,(\hat{\theta}_{n}-\theta_{0})=\frac{A_{n}}{B_{n}-C_{n}}\xrightarrow{% \ \mathrm{d}\ }\frac{N\bigl{(}0,\mathcal{I}(\theta_{0})\bigr{)}}{\mathcal{I}(% \theta_{0})}=N\Bigl{(}0,\frac{1}{\mathcal{I}(\theta_{0})}\Bigr{)},

where in the last step we used the fact that dividing a N⁢(0,σ2)N(0,\sigma^{2}) random variable by the constant c>0c>0 gives a N⁢(0,σ2/c2)N(0,\sigma^{2}/c^{2}) random variable, in this case with σ2=ℐ︀⁢(θ0)\sigma^{2}=\mathcal{I}(\theta_{0}) and c=ℐ︀⁢(θ0)c=\mathcal{I}(\theta_{0}). This completes the proof. ∎

The structure of this argument is typical of most large-sample results in statistics: Taylor expansion is used to transform a question about an implicitly defined estimator into a question about sums of i.i.d. terms, the central limit theorem is used to analyse the leading term, the law of large numbers is used to analyse the coefficient of the leading term, the remainder is shown to be negligible using consistency of the estimator, and Slutsky’s theorem is used to combine the individual results. The only tricky step is the analysis of the remainder term, and condition (R4) is the reason that this term is straightforward to analyse.

17.3 Standard Errors

The variance 1/ℐ︀n⁢(θ0)1/\mathcal{I}_{n}(\theta_{0}) in theorem 17.1 depends on the unknown θ0\theta_{0}, and so cannot be reported as it stands. In practice we replace θ0\theta_{0} by its estimate, and the next result says that this does not disturb the limit.

Corollary 17.3.

With the same assumptions as in theorem 17.1, we have

n⁢ℐ︀⁢(θ^n)⁢(θ^n−θ0)→dN⁢(0,1)(n→∞).\sqrt{n\mathcal{I}(\hat{\theta}_{n})}\,(\hat{\theta}_{n}-\theta_{0})% \xrightarrow{\ \mathrm{d}\ }N(0,1)\qquad(n\to\infty).
Proof.

By (R3) the function ℐ︀\mathcal{I} is continuous and positive on Θ\Theta and thus the function θ↦ℐ︀⁢(θ)/ℐ︀⁢(θ0)\theta\mapsto\sqrt{\mathcal{I}(\theta)/\mathcal{I}(\theta_{0})} is continuous at θ0\theta_{0}. By θ^n→pθ0\hat{\theta}_{n}\xrightarrow{\ \mathrm{p}\ }\theta_{0} from theorem 16.3, the continuous mapping theorem from theorem A.17 implies that ℐ︀⁢(θ^n)/ℐ︀⁢(θ0)→p1\sqrt{\mathcal{I}(\hat{\theta}_{n})/\mathcal{I}(\theta_{0})}\xrightarrow{\ % \mathrm{p}\ }1. By multiplying the result of theorem 17.1 by this factor and using the product part of Slutsky’s theorem, theorem A.16, we find

n⁢ℐ︀⁢(θ^n)⁢(θ^n−θ0)=ℐ︀⁢(θ^n)ℐ︀⁢(θ0)⋅n⁢ℐ︀⁢(θ0)⁢(θ^n−θ0)→d1⋅N⁢(0,1),\sqrt{n\mathcal{I}(\hat{\theta}_{n})}\,(\hat{\theta}_{n}-\theta_{0})=\sqrt{% \frac{\mathcal{I}(\hat{\theta}_{n})}{\mathcal{I}(\theta_{0})}}\cdot\sqrt{n% \mathcal{I}(\theta_{0})}\,(\hat{\theta}_{n}-\theta_{0})\xrightarrow{\ \mathrm{% d}\ }1\cdot N(0,1),

and the claim is proved. ∎

The quantity

se(θ^n)=1ℐ︀n⁢(θ^n)=1n⁢ℐ︀⁢(θ^n)\mathop{\mathrm{se}}\nolimits(\hat{\theta}_{n})=\frac{1}{\sqrt{\mathcal{I}_{n}% (\hat{\theta}_{n})}}=\frac{1}{\sqrt{n\mathcal{I}(\hat{\theta}_{n})}}

is called the standard error of the MLE. As a statistic, this quantity from the corollary is an approximation to the standard deviation of θ^n\hat{\theta}_{n} for large nn. The corollary also allows us to justify approximate confidence statements: Since a standard normal variable takes values in (−1.96,1.96)(-1.96,1.96) with probability 0.950.95, the random interval θ^n±1.96⁢se(θ^n)\hat{\theta}_{n}\pm 1.96\,\mathop{\mathrm{se}}\nolimits(\hat{\theta}_{n}) covers θ0\theta_{0} with probability which tends to 0.950.95. Intervals of this form were used in an informal way in lecture 12 and we will make these intervals precise in lecture 25. In lecture 18 we will use simulation to check how close the real coverage of these intervals is to 0.950.95 for moderate nn. If we do not have a closed formula for ℐ︀\mathcal{I}, the denominator ℐ︀n⁢(θ^n)\mathcal{I}_{n}(\hat{\theta}_{n}) can be replaced by the observed curvature −ℓ′′⁢(θ^n)-\ell^{\prime\prime}(\hat{\theta}_{n}) of the log-likelihood at the maximum. This approach is taken by a numerical optimiser and we will use the resulting values in lecture 18.

17.4 Examples

For the MLEs which are sample means, the Bernoulli and Poisson MLEs from lecture 5, theorem 17.1 does not give any new results: the central limit theorem allows us to find the asymptotic normality of these estimators directly and from exercises 17.1 and 17.2 we know that the two variances are the same. For the normal mean with known variance, the MLE X¯\bar{X} is exactly normally distributed for every nn, with variance σ2/n=1/ℐ︀n⁢(θ)\sigma^{2}/n=1/\mathcal{I}_{n}(\theta). Thus, for this case the theorem gives an exact result. The first new case is an MLE which is not a sample mean.

Example 17.4.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with exponential rate θ\theta. Then the MLE for the parameter value is θ^n=1/X¯\hat{\theta}_{n}=1/\bar{X} from example 5.2 and by exercise 16.1 the regularity conditions and the uniqueness of the root hold. From example 11.8 we know that the Fisher information is given by ℐ︀⁢(θ)=1/θ2\mathcal{I}(\theta)=1/\theta^{2}. Thus, by theorem 17.1,

n⁢(1X¯−θ)→dN⁢(0,θ2),\sqrt{n}\,\Bigl{(}\frac{1}{\bar{X}}-\theta\Bigr{)}\xrightarrow{\ \mathrm{d}\ }% N(0,\theta^{2}),

with standard error se(θ^n)=θ^n/n\mathop{\mathrm{se}}\nolimits(\hat{\theta}_{n})=\hat{\theta}_{n}/\sqrt{n}. We can compare the result to exact results as follows: In lecture 5 we found 𝔼θ⁢(θ^n)=n⁢θ/(n−1)\mathbb{E}_{\theta}(\hat{\theta}_{n})=n\theta/(n-1), which is a bias of θ/(n−1)\theta/(n-1). In exercise 5.3 we found the variance Varθ(θ^n)=n2⁢θ2/((n−1)2⁢(n−2))\mathop{\mathrm{Var}}\nolimits_{\theta}(\hat{\theta}_{n})=n^{2}\theta^{2}/% \bigl{(}(n-1)^{2}(n-2)\bigr{)}. If we multiply this variance by nn, it converges to θ2\theta^{2}. If we multiply the bias by n\sqrt{n}, it converges to zero, as predicted by the theorem: for large nn, the bias is of a smaller order than the spread of the distribution, and thus disappears in the limit.

The theorem can also be applied to functions of the parameter. The mean of the exponential distribution is μ=g⁢(θ)=1/θ\mu=g(\theta)=1/\theta and by theorem 5.6 the MLE for this parameter is μ^n=1/θ^n=X¯\hat{\mu}_{n}=1/\hat{\theta}_{n}=\bar{X}. Using the delta method, theorem A.18 with g′⁢(θ)=−1/θ2g^{\prime}(\theta)=-1/\theta^{2}, we get

n⁢(X¯−μ)→dN⁢(0,g′⁢(θ)2⁢θ2)=N⁢(0,1θ2)=N⁢(0,μ2).\sqrt{n}\,(\bar{X}-\mu)\xrightarrow{\ \mathrm{d}\ }N\bigl{(}0,\ g^{\prime}(% \theta)^{2}\,\theta^{2}\bigr{)}=N\Bigl{(}0,\frac{1}{\theta^{2}}\Bigr{)}=N(0,% \mu^{2}).

This is the central limit theorem for X¯\bar{X}, since the exponential distribution has variance 1/θ2=μ21/\theta^{2}=\mu^{2}. The standard error of X¯\bar{X}, X¯/n\bar{X}/\sqrt{n}, is the standard error we would obtain by computing the standard error from first principles.

The example illustrates a general point: for any smooth function g⁢(θ)g(\theta) of the parameter, theorem 5.6 gives the MLE g⁢(θ^n)g(\hat{\theta}_{n}), and the delta method gives the asymptotic distribution and standard error |g′⁢(θ^n)|⁢se(θ^n)|g^{\prime}(\hat{\theta}_{n})|\,\mathop{\mathrm{se}}\nolimits(\hat{\theta}_{n}), without any further maximisation.

17.5 Quality of the Approximation

Theorem 17.1 is a limit statement and does not give information about the required sample size for the normal approximation to be applicable. The required value of nn for the normal approximation to be usable depends on the model and there are three situations where care is needed. First, for small samples from a model with a skewed distribution: for the exponential rate, θ^n=1/X¯\hat{\theta}_{n}=1/\bar{X} has a right-skewed distribution and a bias of θ/(n−1)\theta/(n-1) which is not captured in the limit but which is visible for nn around 1010. Secondly, when a parameter value is close to the boundary of the parameter space: for the Bernoulli distribution with θ0\theta_{0} close to zero or one, the standard error θ^n⁢(1−θ^n)/n\sqrt{\hat{\theta}_{n}(1-\hat{\theta}_{n})/n} is zero whenever all observations have the same value, and thus the interval θ^n±1.96⁢se(θ^n)\hat{\theta}_{n}\pm 1.96\,\mathop{\mathrm{se}}\nolimits(\hat{\theta}_{n}) has zero width in this case. Finally, if the regularity conditions are violated: for the uniform distribution on (0,θ)(0,\theta), exercise 17.3 shows that the error of the MLE is of order 1/n1/n instead of 1/n1/\sqrt{n}, and that the limiting distribution is exponential instead of normal, as expected from the super-efficiency of the parameter estimates as shown in lecture 13. The first case is investigated by simulation in lecture 12, by checking the coverage of the interval θ^n±1.96⁢se(θ^n)\hat{\theta}_{n}\pm 1.96\,\mathop{\mathrm{se}}\nolimits(\hat{\theta}_{n}) for moderate nn.

Finally, the theorem can be extended to a vector parameter θ=(θ1,…,θk)\theta=(\theta_{1},\dots,\theta_{k}): Under appropriate regularity conditions, the MLE is asymptotically normal with mean θ0\theta_{0} and covariance matrix ℐ︀n⁢(θ0)−1\mathcal{I}_{n}(\theta_{0})^{-1}, the inverse of the Fisher information matrix from lecture 14. Without proof, the argument is similar to the one above, but uses Taylor expansion for several variables. For the normal distribution with unknown parameters, where the information matrix is diagonal, this gives the asymptotic normality of X¯\bar{X} and of the variance estimator separately.

Summary.
  • •
    ​

    Under regularity conditions and assuming the uniqueness of the likelihood root, we have n⁢(θ^n−θ0)→dN⁢(0,1/ℐ︀⁢(θ0))\sqrt{n}\,(\hat{\theta}_{n}-\theta_{0})\xrightarrow{\ \mathrm{d}\ }N\bigl{(}0,% 1/\mathcal{I}(\theta_{0})\bigr{)}: the MLE is asymptotically unbiased, the error is of order 1/n1/\sqrt{n}, and the asymptotic variance equals the Cramer–Rao bound, i.e. the MLE is asymptotically efficient.

  • •
    ​

    The proof uses Taylor expansion of the score around θ0\theta_{0}, central limit theorem for the score, law of large numbers for the second derivative, consistency and (R4) for the remainder, and Slutsky’s theorem to combine these results.

  • •
    ​

    Replacing θ0\theta_{0} by θ^n\hat{\theta}_{n} in the expression for the variance does not change the limit and thus we get the standard error se(θ^n)=1/ℐ︀n⁢(θ^n)\mathop{\mathrm{se}}\nolimits(\hat{\theta}_{n})=1/\sqrt{\mathcal{I}_{n}(\hat{% \theta}_{n})} and the approximate 95%95\% confidence interval θ^n±1.96⁢se(θ^n)\hat{\theta}_{n}\pm 1.96\,\mathop{\mathrm{se}}\nolimits(\hat{\theta}_{n}) for the parameter value.

  • •
    ​

    Using the delta method, we can apply the result to any smooth function of the parameter. The standard error for such an application is |g′⁢(θ^n)|⁢se(θ^n)|g^{\prime}(\hat{\theta}_{n})|\,\mathop{\mathrm{se}}\nolimits(\hat{\theta}_{n}).

  • •
    ​

    The approximation can be poor for small nn in models with skewed distributions and at the boundary of the parameter space, and it breaks down completely when the regularity conditions are violated, e.g. for the uniform distribution.

Exercise 17.1.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Bernoulli with success probability θ∈(0,1)\theta\in(0,1), where the MLE θ^n=X¯\hat{\theta}_{n}=\bar{X} and Fisher information ℐ︀⁢(θ)=1/(θ⁢(1−θ))\mathcal{I}(\theta)=1/\bigl{(}\theta(1-\theta)\bigr{)} are given in example 11.5.

  1. 1.
    ​

    State the result of theorem 17.1 for this model and verify that it coincides with the central limit theorem for X¯\bar{X}.

  2. 2.
    ​

    The log-odds are given by ψ=g⁢(θ)=log⁡(θ/(1−θ))\psi=g(\theta)=\log\bigl{(}\theta/(1-\theta)\bigr{)}. Find the MLE for ψ\psi and use the delta method to show that n⁢(ψ^n−ψ)→dN⁢(0,1/(θ⁢(1−θ)))\sqrt{n}\,(\hat{\psi}_{n}-\psi)\xrightarrow{\ \mathrm{d}\ }N\bigl{(}0,1/(% \theta(1-\theta))\bigr{)}.

  3. 3.
    ​

    Determine the standard error of ψ^n\hat{\psi}_{n} and an approximate 95%95\% interval for ψ\psi, in terms of θ^n\hat{\theta}_{n} and nn.

Exercise 17.2.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Poisson with parameter θ>0\theta>0, where MLE θ^n=X¯\hat{\theta}_{n}=\bar{X} and Fisher information ℐ︀⁢(θ)=1/θ\mathcal{I}(\theta)=1/\theta are given in example 11.6.

  1. 1.
    ​

    State the result of theorem 17.1 for this model and verify that the result coincides with the central limit theorem.

  2. 2.
    ​

    In exercise 5.7 we have found the MLE p^0=e−X¯\hat{p}_{0}=e^{-\bar{X}} for the probability p0=ℙθ⁢(X1=0)=e−θp_{0}=\mathbb{P}_{\theta}(X_{1}=0)=e^{-\theta} and we have seen that this MLE is biased. Using the delta method, determine the asymptotic distribution of n⁢(p^0−p0)\sqrt{n}\,(\hat{p}_{0}-p_{0}). Why does the bias, found in the previous step, not contradict the result from this question?

Exercise 17.3.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. uniformly distributed on the interval (0,θ)(0,\theta). The MLE for this model is θ^n=maxi⁡Xi\hat{\theta}_{n}=\max_{i}X_{i}.

  1. 1.
    ​

    For every x≥0x\geq 0, show that ℙθ⁢(n⁢(θ−θ^n)>x)→e−x/θ\mathbb{P}_{\theta}\bigl{(}n(\theta-\hat{\theta}_{n})>x\bigr{)}\to e^{-x/\theta} as n→∞n\to\infty, i.e. that n⁢(θ−θ^n)n(\theta-\hat{\theta}_{n}) converges in distribution to the exponential distribution with mean θ\theta, i.e. with rate 1/θ1/\theta.

  2. 2.
    ​

    Show that n⁢(θ^n−θ)→p0\sqrt{n}\,(\hat{\theta}_{n}-\theta)\xrightarrow{\ \mathrm{p}\ }0. Why is the MLE not asymptotically normal in the sense of theorem 17.1? Why does this make sense in the context of the super-efficiency of the uniform model from lecture 13?