Lecture 5 Maximum Likelihood Estimation
In the previous lecture we have seen that for the exponential model, the parameter value which maximises the likelihood is the one which ‘best’ explains the data. In this lecture we will make this idea into the main estimation method of the module: we define the maximum likelihood estimator, practise computing it for the families from tables 1.1 and 1.2, and will see problems where more thought is needed. We will see that the estimator is not unbiased in general, but is invariant under reparametrisation; the deeper justification of the method will be given in lectures 13, 16 and 17.
5.1 The Maximum Likelihood Estimator
Example 4.5 already contains the idea: for the exponential sample the log-likelihood increases up to and then decreases, so this value fits the data best. We now formalise this principle.
A maximum likelihood estimator (MLE) for is an estimator , which maximises the likelihood function over the parameter space, i.e. which satisfies
Since the logarithm is strictly increasing, maximises if and only if it maximises the log-likelihood and for the reasons given in lecture 4 we mostly work with . The definition is careful to avoid the problem of a maximum not existing or not being unique. For example, the likelihood could increase as the edge of the parameter space is approached, so that no maximum exists. Similarly, the maximum may not be unique. In such cases, any maximiser can be used. For the standard families of models and typical data, neither of these problems will occur and it will be safe to speak of “the” MLE.
If the set is an interval and the log-likelihood function is differentiable, then calculus can be used to find the MLE: at a maximum inside the interval , the derivative of equals zero and thus the MLE satisfies the likelihood equation
i.e. the MLE is a zero of the score function from definition 4.4. The general procedure is to find the log-likelihood function, to solve the likelihood equation and to verify that the solution is a maximum, for example by checking that or by checking that the score changes from positive to negative.
The solution of the likelihood equation only gives interior critical points. These could be minima or inflection points and the checking step is required to avoid wrong answers. Also, the above procedure assumes that the maximum is interior and that is differentiable. We will see that these assumptions can be violated.
5.2 Worked Examples
We start our analysis by considering the “running families” of the module.
Let be an i.i.d. Bernoulli sample with success probability , and let be the observed number of successes. As in example 4.2 we have the likelihood and thus
Assume first . Then the likelihood equation is
This equation implies and the only solution to this equation is . Since everywhere, this root is the global maximum and the MLE for the proportion of successes is
For the dataset from example 4.2, seven heads in ten tosses, we get , the value which we already recognised in the example.
The assumption is made, because at the two extremes the MLE does not exist: for the function is strictly increasing on and the supremum is approached at the boundary point , which is outside the open parameter space, and thus there is no maximiser. This is a simple case of the existence caveat from definition 5.1. In this case we could, for example, report the boundary value or else extend the parameter space to the closed interval.
Consider , where both parameters are unknown. Then we have . Taking logarithms we get
Since we have a two-dimensional parameter, the likelihood equation translates to a system of two equations: we set both partial derivatives equal to zero. Taking derivatives with respect to we find
which is zero only for . Since the sum of squares takes its minimum for , this value maximises over for every fixed and we can substitute to find the maximum of the remaining function of one variable. Let and . Then we have
with derivative
which is positive for and negative for . Thus, the maximum is at . Hence, the MLE is the pair
This is the plug-in estimator, obtained by using the method of moments in example 4.8, with divisor : for the normal distribution, both methods give the same result for both coordinates. In particular, the MLE for the variance is the biased estimator from section 2.3, where we showed that the bias in this case was harmless. Whether the bias is a general property of maximum likelihood is a question we will consider later.
5.3 When the Recipe Fails
Both assumptions in the recipe, i.e. an interior maximum and differentiable log-likelihood, break down for the uniform model. In this case we have to return to definition 5.1 and maximise the expression directly.
Let be i.i.d., as in example 1.2. From the joint density found in exercise 1.3, the likelihood is
Calculus is useless here: on the region where is positive its derivative is negative everywhere, so the likelihood equation has no solution, and at the function is not even continuous. But maximising directly is easy: the likelihood is zero for and strictly decreasing for , and thus largest at the left edge of the positive region. The MLE is the sample maximum,
This identifies the estimator we have been tracking since exercise 1.3, which by exercise 4.4 beats the method of moments estimator in mean squared error for every . The likelihood also rules out impossible parameter values automatically: for the dataset , , of lecture 4, any has likelihood zero, so the contradiction the method of moments overlooked cannot arise.
The uniform model fails the recipe in an obvious way; a failure which is harder to spot is a likelihood equation with several roots. Consider the location model with degrees of freedom, where is fixed and known and the density is given by
with location and a normalising constant free of . The log-likelihood is differentiable everywhere, so the recipe seems to apply, but after clearing denominators the likelihood equation
turns into a polynomial equation of degree in and for this equation can have several roots. Exercise 5.4 shows that for there are three roots when the observations are more than apart: the mid-point of the interval between the observations, which is then a local minimum of the likelihood, and two global maxima of equal likelihood, one close to each of the observations. Thus, the model prefers to fit one data point and to consider the other as an outlier, and in this case the data cannot choose between the two maxima and the MLE will not be unique. The exercise also shows that a root of the likelihood equation can be a minimum and that the checking step in the recipe is not optional.
For larger samples the likelihood equation of the location model cannot be solved in closed form, and similarly the likelihood equations for many other models, e.g. the two-parameter gamma distribution from example 4.9, cannot be solved in closed form. The solution to this problem, which is to maximise numerically, will be discussed, together with its pitfalls, in lecture 18.
5.4 Bias of the Maximum Likelihood Estimator
Nothing in definition 5.1 mentions bias and indeed the MLE is not always unbiased (we already know this from example 5.4 for the normal distribution). In this section we consider the exponential distribution with rate , where all moments can be determined exactly.
Let be i.i.d. exponential with rate and let be the MLE from example 5.2. Setting , so that , we found in exercise 2.3, using the gamma density of , that for , while for the expectation is infinite. Thus the MLE is biased, with bias , and multiplying by removes the bias: the estimator satisfies .
From lecture 2 we know that removing the bias does not automatically give the estimator with the smaller mean squared error. In exercise 5.3 we compute the second moment of and find for all . Thus, after removing the bias, this estimator is strictly better. The same exercise shows that a third scaling of would be even better, similar to how the divisor in the normal distribution variance was chosen.
For the normal variance the biased divisor- estimator was better than the unbiased sample variance (section 2.3); for the exponential rate the bias-corrected estimator is better. Neither result is a theorem: the bias-variance trade-off can go either way, and the only way to know which is better is to compute both mean squared errors. In favour of the MLE we note that the bias decreases to zero as the sample size increases, and in lecture 17 we will see that for large sample size the MLE is in some sense the best estimator.
5.5 Invariance Under Reparametrisation
Often the quantity we want to estimate is not the parameter itself, but a function of the parameter, e.g. a probability, a mean or a median. The following theorem shows that under maximum likelihood such quantities transform together with the parameter.
Let be a bijection and let be a maximum likelihood estimator for . Then is a maximum likelihood estimator for the transformed parameter .
The idea of the proof is that re-labelling the parameter does not change the family of distributions: in terms of the likelihood is , so maximising over is the same maximisation as before. Writing this down is exercise 5.5.
If is not a bijection, then the likelihood of is defined to be the maximal likelihood of the values of which map to this value. While the statement is still true, we will not require the general case in this module.
As an application, exercise 5.8 finds the MLE for the median waiting time of the exponential model: since is a bijection, theorem 5.6 gives with no further maximisation. The exercise also contains a surprise: is unbiased, despite the MLE for the rate being biased. In general, unbiasedness is lost under nonlinear reparametrisation, and thus there is no analogue of theorem 5.6 for unbiasedness.
Since we have seen that the MLE can be biased, non-unique and even non-existent, we need to do more work to understand the method. The arguments for the validity of the MLE can be found in lectures 13, 16 and 17: the main result is the Cramer–Rao bound, which no unbiased estimator can violate. For large sample sizes, the MLE is shown to be consistent and to asymptotically achieve this bound. In preparation for these arguments, lecture 8 summarises the observation from example 4.3 that the likelihood function can be collapsed to a small number of data summaries.
-
•
The maximum likelihood estimator is the parameter value which maximises the likelihood (or log-likelihood) of the data.
-
•
For smooth problems, where the maximum is an interior point of the parameter space, the MLE can be found by solving the likelihood equation and verifying that the root is a maximum. For the standard families of models this gives (Bernoulli and Poisson), (exponential rate) and (normal).
-
•
If the recipe fails, the definition can still be used: for the uniform distribution the maximum sits at the boundary, at , and for the location model with two observations far apart the MLE is not unique.
-
•
The MLE is not unbiased in general. Whether removing the bias improves the mean squared error depends on the model: for the exponential rate the corrected estimator wins, reversing the normal-variance comparison of lecture 2.
-
•
The MLE is invariant under reparametrisation: the MLE for is . Unbiasedness is not invariant.
- •
Let be an i.i.d. sample from the geometric distribution with parameter , i.e. for as in exercise 4.2.
-
1.
Show that the log-likelihood function is given by .
-
2.
Solve the likelihood equation and show that your solution is a maximum. Deduce that the MLE is .
-
3.
Discuss your results in the light of the method of moments estimator from exercise 4.2.
Let be i.i.d. with density for , where . This is the gamma distribution with known shape and scale , i.e. in the rate parametrisation from table A.2.
-
1.
Find the MLE for .
-
2.
Show that your estimator is unbiased.
-
3.
Determine and use this to show that is consistent.
Let be i.i.d. exponential with rate , let be the MLE and assume . Consider the family of estimators for constants .
-
1.
Using the gamma density of as in exercise 2.3, show that .
-
2.
Show that
and deduce the values for the MLE and for the unbiased version from the text, and the ratio .
-
3.
Determine the value of which minimises the mean squared error, and show that the minimal mean squared error is . Compare your result with the MLE (), the unbiased version , and the result for the normal variance in section 2.3. Comment upon your result.
Let be any two observations from the location model given in section 5.3. Let be the mid-point and be half the distance between the two observations.
-
1.
Let and . Show that the likelihood equation is equivalent to .
-
2.
Show that for the log-likelihood function is , where does not depend on and
-
3.
Show that for the mid-point is the unique MLE, but for the likelihood equation has three roots , , where is a local minimum of and the other two are maximum likelihood estimates with equal likelihood.
Prove the statement of theorem 5.6: Let be a bijection and let the likelihood in the new parametrisation be given by . Show that is the maximum of over .
Let be an observation from the exponential distribution with rate , i.e. for . The following properties of the exponential distribution will be needed later in the module.
-
1.
Show that the exponential distribution has no memory, i.e. that for all we have
-
2.
Show that for all we have
Comment upon your result in the light of the first part of the question.
Let be an i.i.d. sample from the Poisson distribution with parameter , and assume that .
-
1.
In exercise 4.1 we have seen that the score equals zero at . Show that is indeed the MLE for .
-
2.
Determine the MLE for the probability of observing zero.
-
3.
Using the relation for , show that your estimator from the previous part is biased, but asymptotically unbiased. Comment on your result in the light of the fact that is unbiased.
Assume that we have observed the i.i.d. exponential waiting times with rate , and that we want to estimate the median waiting time , i.e. the value such that .
-
1.
Show that and that the map maps to bijectively.
- 2.
-
3.
Show that is an unbiased estimator for , and comment on why this does not contradict the bias of the MLE for the rate observed in this lecture.