Lecture 4 Likelihood and the Method of Moments
In the previous two lectures we have introduced the problem of estimation, but all the estimators we have considered so far were just guesses. In this lecture we introduce the likelihood function, the most important object in this module: the concepts of sufficiency (in lecture 8), Fisher information (in lecture 11), maximum likelihood (in lecture 5) and the Bayesian approach (in lecture 26) are all based on this function. We also discuss the method of moments, a simple method for constructing an estimator from a model. We will see the weaknesses of this approach which will lead to the introduction of the maximum likelihood estimator in the next lecture.
4.1 The Likelihood Function
From lecture 1 we know that a parametric model specifies a joint density for the data, for each . For an i.i.d. sample this is the product of the marginal densities. The idea of this section is to consider the converse problem, i.e. to consider, once the data are fixed, how well a given value of explains the data.
Let be data from a parametric model . Then the likelihood function is the function , given by
and the log-likelihood is , given by
While the formula for is the same as the joint density of the sample, the interpretation of this function has changed: while the joint density is a function of the data with the parameter fixed, the likelihood is a function of the parameter with the data fixed. To emphasise that the data are used to compute the likelihood, we write , and if we substitute the random variables for the numbers , the value of is a random quantity. Both the numerical values and the random quantities can be used freely. For discrete data, equals , i.e. the probability that the model with parameter generates the observed data.
Assume we have observed heads in tosses of a coin with unknown heads probability . Using the Bernoulli weights from table 1.1, we find the likelihood to be
For the fair coin case we have and . Both values are very small, since any sequence of ten tosses is unlikely, but the ratio of these values is more informative: is times better at explaining the data than the fair coin model is, and this is the kind of information we are interested in.
The likelihood is not a probability distribution over : the parameter is an unknown constant, statements like “the probability that is small” make no sense, and does not integrate to one over . We will learn more about this in lecture 26, where we will see how the Bayesian approach can be used to turn the parameter into a random variable.
In practice we almost always use the log-likelihood: the likelihood is a product of factors, typically each less than one, so for realistic sample sizes it is astronomically small. In contrast, is a sum of moderate size and it is much easier to take derivatives of a sum than of a product. Since the logarithm is strictly increasing, the functions and order the parameter values in the same way, so no information is lost when we use the logarithm.
For an i.i.d. sample from the exponential model with density for , the likelihood is
Thus, the log-likelihood is
The product in this expression has been replaced by a smooth function of , which uses only two summaries of the data, and . We will learn more about this approach to summarising data in lecture 8.
4.2 The Score Function
If the log-likelihood function is differentiable, then the derivative is sometimes called the score function. In this section we only use the score function to find the maximum of the likelihood. (We will learn the deeper meaning of the score function, as the carrier of information in the sample, in lecture 11.)
Let be the log-likelihood function and assume that it is differentiable at . Then the score function is given by
The score is the sum of terms, one for each observation, and if we assume that the data are random, these terms are i.i.d. This additive structure allows us to apply the law of large numbers and the central limit theorem to the likelihood function later in the module. Since the score is the derivative of the log-likelihood, a zero of the score is a good starting point to find the parameter value which gives the best explanation for the data.
Continuing from example 4.3, the score for the exponential sample is given by
The score is positive for and negative for , where . Thus the log-likelihood increases up to and then decreases, so is the best explanation for the data. The idea of using the maximiser of the likelihood to estimate parameters is the basis of maximum likelihood estimation, as discussed in lecture 5.
4.3 The Method of Moments
Before we consider how to maximise the likelihood, we will discuss a much simpler method, the oldest general purpose method for constructing estimators. This method is based on the fact that the model often gives the moments of the observations as functions of the model parameters , and that the data can be used to compute the empirical moments. The idea of the method is then to find an estimator such that the two sets of moments coincide.
Let
for . Then is the -th population moment, a deterministic function of , and is the corresponding sample moment, a statistic. By the law of large numbers (theorem A.14 in appendix A), we have : for large sample size the sample moments are close to the population moments corresponding to the true parameter value.
Let be a parameter with components. Then the method of moments estimator is found by solving the system of equations
for , if this system of equations has a unique solution.
The recipe uses as many moment equations as there are unknown parameter components, and the required moments can usually be found in tables like tables A.1 and A.2. We illustrate the method for three of our running examples.
For the exponential model with rate , we have and the only moment equation is . Thus, the method of moments estimator is given by
Coincidentally, this is the same as the zero of the score from example 4.5.
For the normal model with the parameter has two components. Thus we need to match the first two moments, and . We can solve these two moment equations (see exercise 4.3) to find
Note the divisor : since we are matching moments without considering the bias, the method gives the biased plug-in variance estimator from lecture 2, instead of the unbiased sample variance . In this example, from the mean squared error comparison in lecture 2, this is not a problem.
Let be i.i.d. where the shape and the rate are unknown. Then . From table A.2 we know and . Instead of matching and we can equivalently match the mean and the variance: if we write for the plug-in variance, we get the equations
Dividing the first equation by the second we find and thus
This example shows the method in a good light: for the gamma distribution the maximum likelihood equations from lecture 5 have no closed-form solution, whereas the method of moments finds explicit formulas in two lines.
We now weigh the merits of the method. On the credit side, it is quick, often gives explicit formulas where the likelihood does not, and is consistent under weak conditions: the sample moments converge by the law of large numbers, and if the solution of the moment equations depends continuously on the moments, the continuous mapping theorem (theorem A.17) carries the convergence over to the estimator.
The method has three main weaknesses. First, the estimates can be contradicted by the data: for the uniform model we have and thus , but for the dataset , , this gives , an estimate where the observation could never have occurred. This problem is immediately detected by the likelihood, since , but the moment equations do not detect this problem. Secondly, the choice of moments is arbitrary: if we had used the second moment instead of the first, we would in general have obtained a different estimator. Finally, by taking only a few moments, much of the information in the data is lost. A method which makes use of all the information in the data should be more successful. This is the aim of the next few lectures.
In some cases, both approaches coincide: for the exponential model the moment estimator is also the maximiser of the likelihood (from example 4.5), the same is true for the Poisson distribution (exercise 4.1) and for the normal distribution the two methods coincide (lecture 5). In contrast, for the uniform model considered above, the maximum likelihood estimator is instead of and has a very different mean squared error (as shown in exercises 2.2 and 4.4).
4.4 Identifiability
Both recipes of this lecture, and the whole estimation programme, assume that the parameter can be recovered from the distribution of the data: if two parameter values lead to the same distribution, no data can distinguish between them and no good estimator for can exist. This property of a model deserves a name.
A parametric model is identifiable, if different parameter values lead to different distributions, i.e. if we have for all with .
All the families of models listed in tables 1.1 and 1.2 are identifiable. Non-identifiability usually occurs in models which are over-parametrised, as the following example shows.
Assume that we want to model a measurement which is subject to two additive effects. We can assume that are i.i.d. with parameter . Then the distribution of the data depends on only via the sum . Thus, the parameters and give rise to the same distribution and the model is not identifiable. Both recipes from this lecture fail in this case: the likelihood is constant on the lines and the moment equations only contain the unknowns in the combination . To solve the problem we can either reparametrise the model using or we can restrict the parameter space, for example by setting .
Identifiability is a question about the parametrisation, not about the data or the sample size. A more subtle version of the problem is the case of the mixture models from lecture 19, where the models are only identifiable up to relabelling the mixture components. For the rest of the module, we will assume that the models are identifiable, unless we explicitly state otherwise.
-
•
The likelihood is the joint density of the observed data, as a function of with the data fixed, and is not a probability distribution over .
-
•
The score , the derivative of the log-likelihood function , finds the parameter value which best explains the data. The score returns in lecture 11.
-
•
The method of moments solves the equations for . The method is fast, explicit and consistent, but it can return impossible estimates and does not make good use of the data. The maximum likelihood method from lecture 5 is better.
-
•
A model is identifiable, if different parameter values lead to different distributions. If the model is not identifiable, no estimator can work and we need to change the parametrisation.
Let be an i.i.d. sample from a Poisson distribution with parameter and let be the observed data.
-
1.
Write down the likelihood and show that the log-likelihood is .
-
2.
Compute the score and find the value of where the score equals zero.
-
3.
Compute the method of moments estimator for and compare your result to your answer in the previous part.
Let be an i.i.d. sample from the geometric distribution with parameter , i.e. from the distribution with for and . (There is some disagreement among authors whether the geometric distribution counts the number of trials until the first success, inclusive of this trial, starting at as in the present exercise, or the number of failures before the first success, starting at . In this module we always count trials.)
-
1.
Using the method of moments, determine an estimator for .
-
2.
Show that the value of always lies in the interval , i.e. that the estimate can never be contradicted by the data (this is in contrast to the example of the uniform distribution in the text).
Let i.i.d. and let be the method of moments estimator as given in the text.
-
1.
Show that is unbiased and that .
-
2.
In exercise 2.2 we have seen that for the biased estimator . Show that has strictly smaller mean squared error for all and comment on your result.
Consider the model i.i.d. for the parameter .
-
1.
Show that the model is not identifiable.
-
2.
Find a smaller parameter space , where the model is identifiable.
-
3.
On your space , determine the method of moments estimator for , using the second moment.