Lecture 5 Maximum Likelihood Estimation

In the previous lecture we have seen that for the exponential model, the parameter value which maximises the likelihood is the one which ‘best’ explains the data. In this lecture we will make this idea into the main estimation method of the module: we define the maximum likelihood estimator, practise computing it for the families from tables 1.1 and 1.2, and will see problems where more thought is needed. We will see that the estimator is not unbiased in general, but is invariant under reparametrisation; the deeper justification of the method will be given in lectures 13, 16 and 17.

5.1 The Maximum Likelihood Estimator

Example 4.5 already contains the idea: for the exponential sample the log-likelihood increases up to θ=1/x¯\theta=1/\bar{x} and then decreases, so this value fits the data best. We now formalise this principle.

Definition 5.1.

A maximum likelihood estimator (MLE) for θ\theta is an estimator θ^=θ^⁢(X1,…,Xn)\hat{\theta}=\hat{\theta}(X_{1},\dots,X_{n}), which maximises the likelihood function over the parameter space, i.e. which satisfies

L⁢(θ^)≥L⁢(θ)for all θ∈Θ.L(\hat{\theta})\geq L(\theta)\qquad\text{for all $\theta\in\Theta$.}

Since the logarithm is strictly increasing, θ^\hat{\theta} maximises LL if and only if it maximises the log-likelihood ℓ\ell and for the reasons given in lecture 4 we mostly work with ℓ\ell. The definition is careful to avoid the problem of a maximum not existing or not being unique. For example, the likelihood could increase as the edge of the parameter space is approached, so that no maximum exists. Similarly, the maximum may not be unique. In such cases, any maximiser can be used. For the standard families of models and typical data, neither of these problems will occur and it will be safe to speak of “the” MLE.

If the set Θ\Theta is an interval and the log-likelihood function is differentiable, then calculus can be used to find the MLE: at a maximum inside the interval Θ\Theta, the derivative of ℓ\ell equals zero and thus the MLE satisfies the likelihood equation

ℓ′⁢(θ)=0,\ell^{\prime}(\theta)=0,

i.e. the MLE is a zero of the score function from definition 4.4. The general procedure is to find the log-likelihood function, to solve the likelihood equation and to verify that the solution is a maximum, for example by checking that ℓ′′⁢(θ)<0\ell^{\prime\prime}(\theta)<0 or by checking that the score changes from positive to negative.

Warning.

The solution of the likelihood equation only gives interior critical points. These could be minima or inflection points and the checking step is required to avoid wrong answers. Also, the above procedure assumes that the maximum is interior and that ℓ\ell is differentiable. We will see that these assumptions can be violated.

5.2 Worked Examples

We start our analysis by considering the “running families” of the module.

Example 5.2.

In example 4.5, the score of the exponential model equals zero at θ=1/x¯\theta=1/\bar{x}. Since ℓ′′⁢(θ)=−n/θ2<0\ell^{\prime\prime}(\theta)=-n/\theta^{2}<0 for all θ>0\theta>0, the function ℓ\ell is strictly concave and this root is the only global maximum. Thus, the MLE for the rate of the exponential distribution is

θ^=1X¯.\hat{\theta}=\frac{1}{\bar{X}}.

This coincides with the method of moments estimator from example 4.7.

Example 5.3.

Let X1,…,XnX_{1},\dots,X_{n} be an i.i.d. Bernoulli sample with success probability θ\theta, and let s=∑ixis=\sum_{i}x_{i} be the observed number of successes. As in example 4.2 we have the likelihood L⁢(θ)=θs⁢(1−θ)n−sL(\theta)=\theta^{s}(1-\theta)^{n-s} and thus

ℓ⁢(θ)=s⁢log⁡θ+(n−s)⁢log⁡(1−θ),θ∈(0,1).\ell(\theta)=s\log\theta+(n-s)\log(1-\theta),\qquad\theta\in(0,1).

Assume first 0<s<n0<s<n. Then the likelihood equation is

ℓ′⁢(θ)=sθ−n−s1−θ=0.\ell^{\prime}(\theta)=\frac{s}{\theta}-\frac{n-s}{1-\theta}=0.

This equation implies s⁢(1−θ)=(n−s)⁢θs(1-\theta)=(n-s)\theta and the only solution to this equation is θ=s/n\theta=s/n. Since ℓ′′⁢(θ)=−s/θ2−(n−s)/(1−θ)2<0\ell^{\prime\prime}(\theta)=-s/\theta^{2}-(n-s)/(1-\theta)^{2}<0 everywhere, this root is the global maximum and the MLE for the proportion of successes is

θ^=sn=X¯.\hat{\theta}=\frac{s}{n}=\bar{X}.

For the dataset from example 4.2, seven heads in ten tosses, we get θ^=0.7\hat{\theta}=0.7, the value which we already recognised in the example.

Remark.

The assumption 0<s<n0<s<n is made, because at the two extremes the MLE does not exist: for s=ns=n the function ℓ⁢(θ)=n⁢log⁡θ\ell(\theta)=n\log\theta is strictly increasing on (0,1)(0,1) and the supremum is approached at the boundary point θ=1\theta=1, which is outside the open parameter space, and thus there is no maximiser. This is a simple case of the existence caveat from definition 5.1. In this case we could, for example, report the boundary value s/n=1s/n=1 or else extend the parameter space to the closed interval.

Example 5.4.

Consider X1,…,Xn∼N⁢(μ,σ2)X_{1},\dots,X_{n}\sim N(\mu,\sigma^{2}), where both parameters are unknown. Then we have θ=(μ,σ2)\theta=(\mu,\sigma^{2}). Taking logarithms we get

ℓ⁢(μ,σ2)=−n2⁢log⁡(2⁢π⁢σ2)−12⁢σ2⁢∑i=1n(xi−μ)2.\ell(\mu,\sigma^{2})=-\frac{n}{2}\log(2\pi\sigma^{2})-\frac{1}{2\sigma^{2}}% \sum_{i=1}^{n}(x_{i}-\mu)^{2}.

Since we have a two-dimensional parameter, the likelihood equation translates to a system of two equations: we set both partial derivatives equal to zero. Taking derivatives with respect to μ\mu we find

∂ℓ∂μ=1σ2⁢∑i=1n(xi−μ),\frac{\partial\ell}{\partial\mu}=\frac{1}{\sigma^{2}}\sum_{i=1}^{n}(x_{i}-\mu),

which is zero only for μ=x¯\mu=\bar{x}. Since the sum of squares ∑i(xi−μ)2\sum_{i}(x_{i}-\mu)^{2} takes its minimum for μ=x¯\mu=\bar{x}, this value maximises ℓ\ell over μ\mu for every fixed σ2\sigma^{2} and we can substitute μ=x¯\mu=\bar{x} to find the maximum of the remaining function of one variable. Let v=σ2v=\sigma^{2} and W=∑i(xi−x¯)2W=\sum_{i}(x_{i}-\bar{x})^{2}. Then we have

ℓ⁢(x¯,v)=−n2⁢log⁡(2⁢π⁢v)−W2⁢v,\ell(\bar{x},v)=-\frac{n}{2}\log(2\pi v)-\frac{W}{2v},

with derivative

∂ℓ∂v=−n2⁢v+W2⁢v2=W−n⁢v2⁢v2,\frac{\partial\ell}{\partial v}=-\frac{n}{2v}+\frac{W}{2v^{2}}=\frac{W-nv}{2v^% {2}},

which is positive for v<W/nv<W/n and negative for v>W/nv>W/n. Thus, the maximum is at v=W/nv=W/n. Hence, the MLE is the pair

μ^=X¯andσ^2=1n⁢∑i=1n(Xi−X¯)2.\hat{\mu}=\bar{X}\qquad\text{and}\qquad\hat{\sigma}^{2}=\frac{1}{n}\sum_{i=1}^% {n}(X_{i}-\bar{X})^{2}.

This is the plug-in estimator, obtained by using the method of moments in example 4.8, with divisor nn: for the normal distribution, both methods give the same result for both coordinates. In particular, the MLE for the variance is the biased estimator from section 2.3, where we showed that the bias in this case was harmless. Whether the bias is a general property of maximum likelihood is a question we will consider later.

5.3 When the Recipe Fails

Both assumptions in the recipe, i.e. an interior maximum and differentiable log-likelihood, break down for the uniform model. In this case we have to return to definition 5.1 and maximise the expression directly.

Example 5.5.

Let X1,…,Xn∼Uniform⁢(0,θ)X_{1},\dots,X_{n}\sim\text{Uniform}(0,\theta) be i.i.d., as in example 1.2. From the joint density found in exercise 1.3, the likelihood is

L⁢(θ)={θ−nif θ≥maxi⁡xi,0otherwise.L(\theta)=\begin{cases}\theta^{-n}&\text{if $\theta\geq\max_{i}x_{i}$,}\\ 0&\text{otherwise.}\end{cases}

Calculus is useless here: on the region where LL is positive its derivative −n⁢θ−n−1-n\theta^{-n-1} is negative everywhere, so the likelihood equation has no solution, and at θ=maxi⁡xi\theta=\max_{i}x_{i} the function is not even continuous. But maximising directly is easy: the likelihood is zero for θ<maxi⁡xi\theta<\max_{i}x_{i} and strictly decreasing for θ≥maxi⁡xi\theta\geq\max_{i}x_{i}, and thus largest at the left edge of the positive region. The MLE is the sample maximum,

θ^=maxi⁡Xi.\hat{\theta}=\max_{i}X_{i}.

This identifies the estimator we have been tracking since exercise 1.3, which by exercise 4.4 beats the method of moments estimator 2⁢X¯2\bar{X} in mean squared error for every n≥3n\geq 3. The likelihood also rules out impossible parameter values automatically: for the dataset x1=0.1x_{1}=0.1, x2=0.2x_{2}=0.2, x3=0.9x_{3}=0.9 of lecture 4, any θ<0.9\theta<0.9 has likelihood zero, so the contradiction the method of moments overlooked cannot arise.

The uniform model fails the recipe in an obvious way; a failure which is harder to spot is a likelihood equation with several roots. Consider the tt location model with ν≥1\nu\geq 1 degrees of freedom, where ν\nu is fixed and known and the density is given by

f⁢(x;θ)=cν⁢(1+(x−θ)2ν)−(ν+1)/2,x∈ℝ,f(x;\theta)=c_{\nu}\Bigl{(}1+\frac{(x-\theta)^{2}}{\nu}\Bigr{)}^{-(\nu+1)/2},% \qquad x\in\mathbb{R},

with location θ∈ℝ\theta\in\mathbb{R} and a normalising constant cνc_{\nu} free of θ\theta. The log-likelihood is differentiable everywhere, so the recipe seems to apply, but after clearing denominators the likelihood equation

ℓ′⁢(θ)=(ν+1)⁢∑i=1nxi−θν+(xi−θ)2=0\ell^{\prime}(\theta)=(\nu+1)\sum_{i=1}^{n}\frac{x_{i}-\theta}{\nu+(x_{i}-% \theta)^{2}}=0

turns into a polynomial equation of degree 2⁢n−12n-1 in θ\theta and for n≥2n\geq 2 this equation can have several roots. Exercise 5.4 shows that for n=2n=2 there are three roots when the observations are more than 2⁢ν2\sqrt{\nu} apart: the mid-point of the interval between the observations, which is then a local minimum of the likelihood, and two global maxima of equal likelihood, one close to each of the observations. Thus, the model prefers to fit one data point and to consider the other as an outlier, and in this case the data cannot choose between the two maxima and the MLE will not be unique. The exercise also shows that a root of the likelihood equation can be a minimum and that the checking step in the recipe is not optional.

For larger samples the likelihood equation of the tt location model cannot be solved in closed form, and similarly the likelihood equations for many other models, e.g. the two-parameter gamma distribution from example 4.9, cannot be solved in closed form. The solution to this problem, which is to maximise ℓ\ell numerically, will be discussed, together with its pitfalls, in lecture 18.

5.4 Bias of the Maximum Likelihood Estimator

Nothing in definition 5.1 mentions bias and indeed the MLE is not always unbiased (we already know this from example 5.4 for the normal distribution). In this section we consider the exponential distribution with rate θ\theta, where all moments can be determined exactly.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. exponential with rate θ\theta and let θ^=1/X¯\hat{\theta}=1/\bar{X} be the MLE from example 5.2. Setting T=X1+⋯+XnT=X_{1}+\dots+X_{n}, so that θ^=n/T\hat{\theta}=n/T, we found in exercise 2.3, using the gamma density of TT, that 𝔼θ⁢(θ^)=n⁢θ/(n−1)\mathbb{E}_{\theta}(\hat{\theta})=n\theta/(n-1) for n≥2n\geq 2, while for n=1n=1 the expectation is infinite. Thus the MLE is biased, with bias θ/(n−1)\theta/(n-1), and multiplying by (n−1)/n(n-1)/n removes the bias: the estimator θ~=(n−1)⁢θ^/n=(n−1)/T\tilde{\theta}=(n-1)\hat{\theta}/n=(n-1)/T satisfies 𝔼θ⁢(θ~)=θ\mathbb{E}_{\theta}(\tilde{\theta})=\theta.

From lecture 2 we know that removing the bias does not automatically give the estimator with the smaller mean squared error. In exercise 5.3 we compute the second moment of θ^\hat{\theta} and find MSE(θ~)/MSE(θ^)=(n−1)/(n+2)<1\mathop{\mathrm{MSE}}\nolimits(\tilde{\theta})/\mathop{\mathrm{MSE}}\nolimits(% \hat{\theta})=(n-1)/(n+2)<1 for all n≥3n\geq 3. Thus, after removing the bias, this estimator is strictly better. The same exercise shows that a third scaling of θ^\hat{\theta} would be even better, similar to how the divisor n+1n+1 in the normal distribution variance was chosen.

For the normal variance the biased divisor-nn estimator was better than the unbiased sample variance S2S^{2} (section 2.3); for the exponential rate the bias-corrected estimator is better. Neither result is a theorem: the bias-variance trade-off can go either way, and the only way to know which is better is to compute both mean squared errors. In favour of the MLE we note that the bias decreases to zero as the sample size increases, and in lecture 17 we will see that for large sample size the MLE is in some sense the best estimator.

5.5 Invariance Under Reparametrisation

Often the quantity we want to estimate is not the parameter itself, but a function of the parameter, e.g. a probability, a mean or a median. The following theorem shows that under maximum likelihood such quantities transform together with the parameter.

Theorem 5.6 (Invariance of the MLE).

Let g:Θ→Ψg\colon\Theta\to\Psi be a bijection and let θ^\hat{\theta} be a maximum likelihood estimator for θ\theta. Then g⁢(θ^)g(\hat{\theta}) is a maximum likelihood estimator for the transformed parameter ψ=g⁢(θ)\psi=g(\theta).

The idea of the proof is that re-labelling the parameter does not change the family of distributions: in terms of ψ\psi the likelihood is L~⁢(ψ)=L⁢(g−1⁢(ψ))\tilde{L}(\psi)=L\bigl{(}g^{-1}(\psi)\bigr{)}, so maximising L~\tilde{L} over Ψ\Psi is the same maximisation as before. Writing this down is exercise 5.5.

Remark.

If gg is not a bijection, then the likelihood of ψ\psi is defined to be the maximal likelihood of the values of θ\theta which map to this value. While the statement ψ^=g⁢(θ^)\hat{\psi}=g(\hat{\theta}) is still true, we will not require the general case in this module.

As an application, exercise 5.8 finds the MLE for the median waiting time m=(log⁡2)/θm=(\log 2)/\theta of the exponential model: since θ↦(log⁡2)/θ\theta\mapsto(\log 2)/\theta is a bijection, theorem 5.6 gives m^=(log⁡2)⁢X¯\hat{m}=(\log 2)\,\bar{X} with no further maximisation. The exercise also contains a surprise: m^\hat{m} is unbiased, despite the MLE for the rate being biased. In general, unbiasedness is lost under nonlinear reparametrisation, and thus there is no analogue of theorem 5.6 for unbiasedness.

Since we have seen that the MLE can be biased, non-unique and even non-existent, we need to do more work to understand the method. The arguments for the validity of the MLE can be found in lectures 13, 16 and 17: the main result is the Cramer–Rao bound, which no unbiased estimator can violate. For large sample sizes, the MLE is shown to be consistent and to asymptotically achieve this bound. In preparation for these arguments, lecture 8 summarises the observation from example 4.3 that the likelihood function can be collapsed to a small number of data summaries.

Summary.
  • •
    ​

    The maximum likelihood estimator θ^\hat{\theta} is the parameter value which maximises the likelihood (or log-likelihood) of the data.

  • •
    ​

    For smooth problems, where the maximum is an interior point of the parameter space, the MLE can be found by solving the likelihood equation ℓ′⁢(θ)=0\ell^{\prime}(\theta)=0 and verifying that the root is a maximum. For the standard families of models this gives X¯\bar{X} (Bernoulli and Poisson), 1/X¯1/\bar{X} (exponential rate) and (X¯,(1/n)⁢∑i(Xi−X¯)2)\bigl{(}\bar{X},(1/n)\sum_{i}(X_{i}-\bar{X})^{2}\bigr{)} (normal).

  • •
    ​

    If the recipe fails, the definition can still be used: for the uniform distribution the maximum sits at the boundary, at θ^=maxi⁡Xi\hat{\theta}=\max_{i}X_{i}, and for the tt location model with two observations far apart the MLE is not unique.

  • •
    ​

    The MLE is not unbiased in general. Whether removing the bias improves the mean squared error depends on the model: for the exponential rate the corrected estimator wins, reversing the normal-variance comparison of lecture 2.

  • •
    ​

    The MLE is invariant under reparametrisation: the MLE for g⁢(θ)g(\theta) is g⁢(θ^)g(\hat{\theta}). Unbiasedness is not invariant.

  • •
    ​

    The full justification of the method, including efficiency and large-sample optimality, follows in lectures 13, 16 and 17.

Exercise 5.1.

Let X1,…,XnX_{1},\dots,X_{n} be an i.i.d. sample from the geometric distribution with parameter p∈(0,1)p\in(0,1), i.e. f⁢(x;p)=(1−p)x−1⁢pf(x;p)=(1-p)^{x-1}p for x∈{1,2,…}x\in\{1,2,\dots\} as in exercise 4.2.

  1. 1.
    ​

    Show that the log-likelihood function is given by ℓ⁢(p)=n⁢log⁡p+(∑ixi−n)⁢log⁡(1−p)\ell(p)=n\log p+\bigl{(}\sum_{i}x_{i}-n\bigr{)}\log(1-p).

  2. 2.
    ​

    Solve the likelihood equation and show that your solution is a maximum. Deduce that the MLE is p^=1/X¯\hat{p}=1/\bar{X}.

  3. 3.
    ​

    Discuss your results in the light of the method of moments estimator from exercise 4.2.

Exercise 5.2.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with density f⁢(x;θ)=θ−2⁢x⁢e−x/θf(x;\theta)=\theta^{-2}x\,e^{-x/\theta} for x≥0x\geq 0, where θ>0\theta>0. This is the gamma distribution with known shape 22 and scale θ\theta, i.e. Gamma⁢(2,1/θ)\text{Gamma}(2,1/\theta) in the rate parametrisation from table A.2.

  1. 1.
    ​

    Find the MLE θ^\hat{\theta} for θ\theta.

  2. 2.
    ​

    Show that your estimator θ^\hat{\theta} is unbiased.

  3. 3.
    ​

    Determine MSE(θ^)\mathop{\mathrm{MSE}}\nolimits(\hat{\theta}) and use this to show that θ^\hat{\theta} is consistent.

Exercise 5.3.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. exponential with rate θ\theta, let θ^=1/X¯\hat{\theta}=1/\bar{X} be the MLE and assume n≥3n\geq 3. Consider the family of estimators c⁢θ^c\hat{\theta} for constants c>0c>0.

  1. 1.
    ​

    Using the gamma density of T=X1+⋯+XnT=X_{1}+\dots+X_{n} as in exercise 2.3, show that 𝔼θ⁢(θ^2)=n2⁢θ2/((n−1)⁢(n−2))\mathbb{E}_{\theta}(\hat{\theta}^{2})=n^{2}\theta^{2}/\bigl{(}(n-1)(n-2)\bigr{)}.

  2. 2.
    ​

    Show that

    MSE(c⁢θ^)=(n2(n−1)⁢(n−2)⁢c2−2⁢nn−1⁢c+1)⁢θ2,\mathop{\mathrm{MSE}}\nolimits(c\hat{\theta})=\Bigl{(}\frac{n^{2}}{(n-1)(n-2)}% \,c^{2}-\frac{2n}{n-1}\,c+1\Bigr{)}\theta^{2},

    and deduce the values MSE(θ^)=(n+2)⁢θ2/((n−1)⁢(n−2))\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})=(n+2)\theta^{2}/\bigl{(}(n-1)(n-2% )\bigr{)} for the MLE and MSE(θ~)=θ2/(n−2)\mathop{\mathrm{MSE}}\nolimits(\tilde{\theta})=\theta^{2}/(n-2) for the unbiased version θ~=(n−1)⁢θ^/n\tilde{\theta}=(n-1)\hat{\theta}/n from the text, and the ratio MSE(θ~)/MSE(θ^)=(n−1)/(n+2)\mathop{\mathrm{MSE}}\nolimits(\tilde{\theta})/\mathop{\mathrm{MSE}}\nolimits(% \hat{\theta})=(n-1)/(n+2).

  3. 3.
    ​

    Determine the value of cc which minimises the mean squared error, and show that the minimal mean squared error is θ2/(n−1)\theta^{2}/(n-1). Compare your result with the MLE (c=1c=1), the unbiased version c=(n−1)/nc=(n-1)/n, and the result for the normal variance in section 2.3. Comment upon your result.

Exercise 5.4.

Let x1>x2x_{1}>x_{2} be any two observations from the tt location model given in section 5.3. Let m=(x1+x2)/2m=(x_{1}+x_{2})/2 be the mid-point and d=(x1−x2)/2d=(x_{1}-x_{2})/2 be half the distance between the two observations.

  1. 1.
    ​

    Let a=x1−θa=x_{1}-\theta and b=x2−θb=x_{2}-\theta. Show that the likelihood equation ℓ′⁢(θ)=0\ell^{\prime}(\theta)=0 is equivalent to (a+b)⁢(ν+a⁢b)=0(a+b)(\nu+ab)=0.

  2. 2.
    ​

    Show that for θ=m+t\theta=m+t the log-likelihood function is ℓ⁢(m+t)=C−ν+12⁢log⁡P⁢(t)\ell(m+t)=C-\frac{\nu+1}{2}\log P(t), where CC does not depend on tt and

    P⁢(t)=(ν+d2)2+2⁢(ν−d2)⁢t2+t4.P(t)=(\nu+d^{2})^{2}+2(\nu-d^{2})\,t^{2}+t^{4}.
  3. 3.
    ​

    Show that for d2<νd^{2}<\nu the mid-point mm is the unique MLE, but for d2>νd^{2}>\nu the likelihood equation has three roots mm, m±d2−νm\pm\sqrt{d^{2}-\nu}, where mm is a local minimum of ℓ\ell and the other two are maximum likelihood estimates with equal likelihood.

Exercise 5.5.

Prove the statement of theorem 5.6: Let g:Θ→Ψg\colon\Theta\to\Psi be a bijection and let the likelihood in the new parametrisation be given by L~⁢(ψ)=L⁢(g−1⁢(ψ))\tilde{L}(\psi)=L\bigl{(}g^{-1}(\psi)\bigr{)}. Show that g⁢(θ^)g(\hat{\theta}) is the maximum of L~\tilde{L} over Ψ\Psi.

Exercise 5.6.

Let XX be an observation from the exponential distribution with rate θ>0\theta>0, i.e. ℙθ⁢(X>x)=e−θ⁢x\mathbb{P}_{\theta}(X>x)=e^{-\theta x} for x≥0x\geq 0. The following properties of the exponential distribution will be needed later in the module.

  1. 1.
    ​

    Show that the exponential distribution has no memory, i.e. that for all s,t≥0s,t\geq 0 we have

    ℙθ⁢(X>s+t⁢|X>⁢s)=ℙθ⁢(X>t).\mathbb{P}_{\theta}(X>s+t\mskip 1.0mu|\mskip 1.0muX>s)=\mathbb{P}_{\theta}(X>t).
  2. 2.
    ​

    Show that for all c≥0c\geq 0 we have

    𝔼θ⁢(X⁢|X>⁢c)=c+1θ.\mathbb{E}_{\theta}(X\mskip 1.0mu|\mskip 1.0muX>c)=c+\frac{1}{\theta}.

    Comment upon your result in the light of the first part of the question.

Exercise 5.7.

Let X1,…,XnX_{1},\dots,X_{n} be an i.i.d. sample from the Poisson distribution with parameter θ>0\theta>0, and assume that ∑ixi≥1\sum_{i}x_{i}\geq 1.

  1. 1.
    ​

    In exercise 4.1 we have seen that the score equals zero at x¯\bar{x}. Show that X¯\bar{X} is indeed the MLE for θ\theta.

  2. 2.
    ​

    Determine the MLE for the probability p0=ℙθ⁢(X1=0)p_{0}=\mathbb{P}_{\theta}(X_{1}=0) of observing zero.

  3. 3.
    ​

    Using the relation 𝔼θ⁢(zX1)=eθ⁢(z−1)\mathbb{E}_{\theta}\bigl{(}z^{X_{1}}\bigr{)}=e^{\theta(z-1)} for z>0z>0, show that your estimator from the previous part is biased, but asymptotically unbiased. Comment on your result in the light of the fact that X¯\bar{X} is unbiased.

Exercise 5.8.

Assume that we have observed the i.i.d. exponential waiting times X1,…,XnX_{1},\dots,X_{n} with rate θ>0\theta>0, and that we want to estimate the median waiting time mm, i.e. the value such that ℙθ⁢(X1≤m)=1/2\mathbb{P}_{\theta}(X_{1}\leq m)=1/2.

  1. 1.
    ​

    Show that m=g⁢(θ)=(log⁡2)/θm=g(\theta)=(\log 2)/\theta and that the map gg maps (0,∞)(0,\infty) to (0,∞)(0,\infty) bijectively.

  2. 2.
    ​

    Using theorem 5.6 and example 5.2, show that the MLE for the median is given by m^=(log⁡2)⁢X¯\hat{m}=(\log 2)\,\bar{X}.

  3. 3.
    ​

    Show that m^\hat{m} is an unbiased estimator for mm, and comment on why this does not contradict the bias of the MLE θ^=1/X¯\hat{\theta}=1/\bar{X} for the rate observed in this lecture.