Lecture 11 Fisher Information
Since the data are random, the score function from lecture 4 is a random variable, and the Fisher information, given by the variance of the score, measures how much information the sample contains about the parameter. In lecture 13 we will use the Fisher information to obtain a lower bound for the variance of any unbiased estimator, and in lecture 17 we will see that the Fisher information equals the asymptotic variance of the maximum likelihood estimator. In this lecture we will define the Fisher information, prove two identities which allow us to compute the Fisher information, and we will compute the Fisher information for the standard families of models.
11.1 Setting and Assumptions
Throughout this lecture, are i.i.d. with density or probability weights , where the parameter space is an open interval. As in lecture 4, the log-likelihood is the sum of the terms and the score of definition 4.4 is its derivative . In this section we study the distribution of the score under , for a fixed value of .
The results of this lecture are obtained by differentiating, with respect to , the identity , and we need this differentiation to be legitimate. We assume the following throughout. The support of the distribution does not depend on . For every the function is twice continuously differentiable on . Finally, the integral of over the support may be differentiated under the integral sign twice, so that
with sums in place of integrals for a discrete model. These three assumptions are referred to as the conditions of this section. Most standard families of models from tables 1.1 and 1.2 satisfy these conditions, except for the uniform distribution (in section 11.7). The exponential families in the sense of definition 7.1 with twice continuously differentiable natural parameter also satisfy these conditions (since the exchange is the step in the proof of theorem 7.7). These conditions are included in the full list of regularity conditions in definition 16.1 of lecture 16.
11.2 The Score Has Mean Zero
The first identity is that, if is the correct parameter value, the data do not systematically deviate from this value: the score at the true parameter value is zero on average.
Using the chain rule, we find the derivative of as
Since the denominator is positive (because is in the support), we can cancel the density with the denominator to find
where the integral is over the support of the distribution, which does not depend on , and the third equality comes from the case of (11.1). Summing over gives the statement for the whole sample. This completes the proof. ∎
If the parameter value is wrong, the score is not mean zero: for the exponential model from example 4.5, the score has expectation if the true parameter value is , positive for and negative for . The score on average points towards the correct parameter value from both sides. This is the mechanism behind the consistency proof of lecture 16.
11.3 Fisher Information
The next question we could ask is about the fluctuations of the score around zero. If the score is typically large in magnitude, this indicates that the data allow us to distinguish well between nearby parameter values. In contrast, if the score is typically close to zero, this indicates that the log-likelihood is nearly flat at the truth. The variance of the score is a measure for how informative the sample is.
Under the conditions of section 11.1, the Fisher information of the sample about is given by
where the variance of the score is taken at . The Fisher information of a single observation is denoted by and is given by the variance of .
The variance of the score is often awkward to compute directly, and the second expression of the following theorem is nearly always the quickest route to the information.
The first equality follows from the fact that has mean zero by lemma 11.1, and thus the variance equals the second moment. The second equality is obtained by taking the derivative of the ratio (11.2) using the quotient rule:
Similar to the proof of lemma 11.1, the expectation of the first term cancels out in this case, using the case of (11.1):
Thus we have for all . The terms are independent of each other, since they are functions of different observations, and since they have mean zero, the variance of the sum is the sum of the variances:
where we used in the last step. This completes the proof. ∎
The second expression gives the name of the quantity: it is the expected curvature of the log-likelihood at the true parameter value. Large information means that the log-likelihood has a sharp maximum, well-defined by the data. Small information means the log-likelihood is flat. In lecture 17 this idea will be turned into a theorem: the approximate variance of the maximum likelihood estimator is for large sample size, and the numerical value for the curvature at the maximum likelihood estimate can be obtained by a numerical optimiser (lecture 18).
The proof also shows how the information of the sample relates to the information of one observation.
The score is the sum of independent terms and thus the variance can be found as the sum , as in the proof of theorem 11.3, and since the observations are identically distributed, each term in the sum equals
This completes the proof. ∎
11.4 Examples
We now compute the information for the standard families of tables 1.1 and 1.2. By proposition 11.4 it suffices to work with a single observation .
For the Bernoulli distribution with success probability we have and thus
Using we find
Thus, the information of the sample is . Exercise 11.1 checks these identities for this model directly.
For the Poisson distribution with mean we have and thus
In the case of the normal distribution with unknown mean and known variance , the parameter is and . Thus we have
Since the second derivative does not depend on the data, we do not need to consider the expectation. Note that by lemma 2.2 and from lecture 13 we know that no unbiased estimator can be better.
For the exponential distribution with rate , we have and thus
Again, the second derivative is non-random. In fact, section 11.6 shows that the dependence of the information on the scale of the parameter is a general property.
11.5 Exponential Families
All four models from section 11.4 are exponential families. For such models, the information is the second derivative of the cumulant function , given by definition 7.6.
Let be an exponential family in its natural parametrisation. Then the Fisher information of a single observation about is given by
On the other hand, if the family is written as as in definition 7.1 and if is two times continuously differentiable, then the information about is given by
In the natural parametrisation we have and thus and . By theorem 7.7 we have and and thus, since the second derivative is non-random, we find . For the general parametrisation we have , since both quantities are determined by the requirement that the density must integrate to one, and thus, using the chain rule, we find
The factor is a constant and thus the variance of the score is
as claimed. ∎
11.6 Reparametrisation
The information depends on which parameter we use to describe a model, and the factor in proposition 11.9 is an instance of the general rule.
Let where is a differentiable bijection with for all . Let denote the Fisher information for a single observation when the model is parametrised by and let be the information when the model is parametrised by . Then,
With the new parametrisation, the density is and the log-likelihood for a single observation is . Since never vanishes on the interval , by Darboux’s theorem is strictly monotonic and is differentiable. Using the chain rule and the derivative of the inverse function we get for . This gives
Since is fixed for fixed , is a constant and taking variances under gives , as claimed. ∎
For the exponential distribution, reparametrising from the rate to the mean changes the information from example 11.8 to ; see exercise 11.5. Thus, the information is not a property of the model in isolation, but depends on the parametrisation. What does not change under is the combination , since , and this invariance is the basis of the Jeffreys prior in lecture 26.
11.7 When the Identities Fail
The uniform distribution from example 1.2 shows what can go wrong if the conditions from section 11.1 are violated.
Let be uniformly distributed on , i.e. for . The support of the distribution depends on , but we can still compute the derivatives on the support to see whether the identities hold. For we have and thus and . The score is a negative constant and thus we have , in contradiction to lemma 11.1. The variance is zero, whereas and : the three expressions for theorem 11.3 take three different values, one of which is negative. The proof breaks down at the interchange between differentiation and integration: for this model, the two sides of (11.1) with differ by the term from the moving upper limit, which is missed by the differentiation under the integral sign. Exercise 11.6 can be used to work out the two computations.
For this model, there is no Fisher information and the two results based on it are not applicable: the sample maximum does not satisfy the lower bound from lecture 13, the mean squared error decreases like per exercise 2.2 instead of like , and the asymptotic normality from lecture 17 also fails.
-
•
Under the conditions of section 11.1 (a support free of , twice continuously differentiable density and differentiation under the integral sign), the score at the true parameter value has mean zero.
-
•
The Fisher information is the variance of the score, and it equals and , the expected curvature of the log-likelihood at the truth.
-
•
Information from independent observations can be combined, ; for an exponential family we have ; under a reparametrisation we get .
-
•
We have learned the following standard values: for the Bernoulli distribution, for the Poisson distribution, for the mean of the normal distribution and for the rate parameter in the exponential distribution. For the uniform distribution these relations do not hold, since the support of the distribution moves with the parameter.
- •
Let be i.i.d. Bernoulli with success probability , and let be the score of a single observation from example 11.5.
Let be i.i.d. with density for , where , the gamma distribution with known shape and scale from exercise 5.2.
-
1.
Compute the score of a single observation and verify that it has mean zero.
-
2.
Show that , once via and once via the variance of the score, using .
- 3.
-
4.
In exercise 5.2 we found that the MLE is unbiased with . Compare this with .
Let be i.i.d. from the geometric distribution with parameter , i.e. for , where and , as in exercises 4.2 and 5.1.
-
1.
Compute the score of a single observation and verify that it has mean zero.
-
2.
Determine and write down .
-
3.
Verify the value from the previous part by taking the variance of the score.
Let be i.i.d. Poisson with mean .