Lecture 14 Multiparameter Fisher Information
In lectures 11 and 13 the parameter was a single number, but most models we actually fit, e.g. the normal distribution with unknown mean and variance, have several parameters. In this lecture we will extend the concepts of Fisher information and the Cramer–Rao bound to a vector parameter : the score will be a vector, the information will be a matrix, and the bound will be an inequality of covariance matrices. The new feature here is that we will need to invert the matrix, and the inverse will give the cost of estimating one parameter in the presence of the other parameters.
14.1 The Score Vector and the Information Matrix
Setting the stage as in lecture 11, we now assume that the parameter space is an open set and that the derivatives with respect to are partial derivatives, denoted by for the derivative with respect to . The data are now i.i.d. with density or probability weights and the log-likelihood function is , where , and we assume that the conditions of section 11.1 hold for each coordinate, allowing for differentiation under the integral sign for any pair of coordinates, and we assume that each coordinate of the score has finite variance under , so that the entries of the information matrix below are finite.
The score vector of the sample is given by the gradient of the log-likelihood function
as a column vector in .
For the score vector is the score from definition 4.4. Setting the score vector to zero leads to the system of likelihood equations for the normal distribution considered in example 5.4. As in lecture 11, we will consider the score vector to be a random vector under and we will determine the covariance of this random vector. The covariance matrix of a random vector has entries . The diagonal of this matrix contains the variances.
Given the conditions of section 11.1, the Fisher information matrix of the sample about is defined as
where this is the covariance matrix of the score vector. The information matrix for a single observation is denoted by .
The diagonal element is the Fisher information about in the sense of definition 11.2 with all other coordinates fixed to their true values; the off-diagonal elements are new and measure the strength of correlations between the components of the score. Both identities from lecture 11 carry over to this situation, separately for each coordinate.
With the assumptions of section 11.1 in place, we have and
Thus, is minus the expected Hessian matrix of the log-likelihood.
The proof of this result is similar to the proof of lemma 11.1 and theorem 11.3, but with partial derivatives instead of . We omit the details of the proof here, but we can mention that the ratio has mean zero since the density in the numerator and denominator cancels. Using the quotient rule we find where the first term has mean zero. Finally, the independence of the observations allows us to carry over the resulting identity from one observation to the sample. As for the scalar case, the second expression is the more practical one, at least for , because it requires only three second derivatives (one mixed derivative, occurring twice by symmetry).
14.2 Properties of the Information Matrix
The three properties of the scalar information which we used most often in lecture 11 have matrix counterparts. Recall that a symmetric matrix is positive semi-definite, if for all , and positive definite, if for all ; a positive definite matrix has a positive definite inverse.
Suppose the conditions from section 11.1 hold. Then the following statements hold.
-
1.
The matrix is symmetric and positive semi-definite. In particular we have for all .
-
2.
Information is additive: .
-
3.
Let be a continuously differentiable bijection between open subsets of — here is a parameter vector, the several-parameter counterpart of the in lecture 5 — and let be the invertible Jacobian matrix with entries . Then the information matrix for is given by
For the first statement, the symmetry of the matrix follows from the fact that for all and . For , the variance of the linear combination is given by
and thus is non-negative. The second statement is proved by using the fact that the score vector of the sample is the sum of i.i.d. score vectors of the individual observations and that the covariance matrix of the sum of independent random vectors equals the sum of the covariance matrices. Finally, for the third statement, we note that the log-likelihood in the new parametrisation is and thus we find
i.e. . Since is a fixed matrix for fixed , the covariance matrix of is given by . This completes the proof. ∎
For the third statement is a consequence of proposition 11.10: If we have and . Positive definiteness is violated in cases where the log-likelihood does not respond to changes of in some direction, for example when the model is not identifiable in the sense of definition 4.10. From now on we assume that is positive definite so that the inverse exists.
14.3 The Cramer–Rao Bound for Vector Parameters
Theorem 13.2 bounds the variance of an unbiased estimator for a scalar parameter from below. For a vector parameter the natural object to bound is the covariance matrix, and the natural reading of “at least” between symmetric matrices is that the difference is positive semi-definite.
Suppose we have the conditions of section 11.1 and a positive definite matrix . Let be an unbiased estimator for , so that for all and all , where the variances are finite. Then the matrix
is positive semi-definite for all . In particular we have
for and for all .
We only sketch the argument, which is the proof of theorem 13.2 for linear combinations. Differentiating under the integral sign, as in that proof, gives for and otherwise, and thus by bilinearity for all . Using the Cauchy–Schwarz inequality, lemma A.8, and the first statement of proposition 14.4, we find
Choosing , both and equal , which is positive for , and dividing by this value gives . This is the claim, since , and for the diagonal case we have , with the th unit vector. This completes the sketch of the proof. ∎
The bound is almost always applied in its diagonal form: The variance of any unbiased estimator for , when the other components are unknown, is at least the th diagonal entry of the inverse information matrix. Note that this is the diagonal of the inverse, not the reciprocal of the diagonal; we will see the difference in section 14.5. An estimator which achieves the bound for is said to be efficient for , as in lecture 13. The bound for a function with gradient is , which includes the statement of theorem 13.3 for , as well as the linear case , which the theorem already covers. The result is also used in exercise 14.3.
14.4 The Normal Distribution
The only example of a distribution which all students should be able to reproduce without notes is the normal distribution with unknown parameters. We will consider this example in detail.
Let be i.i.d. where , both unknown. From example 5.4 we know that the log-likelihood function is
and the first partial derivatives are given by
Taking second derivatives, where we treat as a single variable, we find the three second derivatives as
To compute the expectations, we use the fact that and . The mixed derivative has expectation zero. The last term has expectation . Thus, by proposition 14.3, we find the information matrix and its inverse to be
We should commit these matrices to memory. The upper left entry is the information about the mean, where the variance is known from example 11.7. The sample mean, with variance , achieves the bound for . For the variance the bound is . The UMVUE from example 10.7 has variance , as shown in the computation of section 2.3, and thus does not achieve the bound. Since is the UMVUE, no unbiased estimator for achieves the bound when is unknown. If we had parametrised the model by the standard deviation instead, i.e. if we had , then the third statement of proposition 14.4, with , gives . See exercise 14.1 for details.
The information matrix for the normal distribution is diagonal. The following section explains the significance of this fact and what happens when this is not the case.
14.5 Nuisance Parameters
Assume that we are interested in only, and that the remaining components are nuisance parameters, i.e. unknown but not of interest. If the values of the nuisance parameters were known, the Cramer–Rao bound for would be . Since the values are unknown, the bound is by theorem 14.5. For we can compare the two bounds explicitly. For this we write
where since the matrix is positive definite, and
is the correlation between the two components of the score vector. Since , the bound where the nuisance parameter is unknown is times larger than the bound where the nuisance parameter is known, and both bounds are equal if and only if . Two parameters with are said to be orthogonal; in this case, not knowing does not incur any additional cost for the bound when estimating . The same inequality holds for all and all : not knowing a parameter can never be beneficial.
The normal distribution is the orthogonal case: in example 14.6 the mixed derivative has mean zero, and the bound for is whether or not is known. This is a special feature of the normal distribution, not the rule. For the gamma distribution with shape and rate both unknown, as in exercise 7.3, all three second derivatives of the log-likelihood of one observation are non-random, and the information matrix of one observation is
where , the second derivative of , is the trigamma function; exercise 14.2 derives this matrix. The off-diagonal entry is non-zero and by (14.2) we have ; for , with , this gives , so that the bound with the other parameter unknown is about four and a half times the bound with the other parameter known.
The inverse information matrix is needed when we estimate several parameters simultaneously. In lecture 17 we will see that, for large sample size, the maximum likelihood estimator is approximately normally distributed with covariance matrix . In lecture 18 we will see how we can numerically find the standard errors of the estimator by inverting the observed Hessian matrix, and in lecture 23 we will see that the number of parameters can be seen as the number of degrees of freedom in this system.
-
•
For a vector parameter the score is the gradient , and the Fisher information matrix is its covariance matrix, equal to minus the expected Hessian of the log-likelihood.
-
•
The information matrix is symmetric and positive semi-definite, adds over independent observations, , and transforms as under a change of parametrisation with Jacobian .
-
•
Cramer–Rao: for an unbiased estimator for the matrix is positive semi-definite; in particular .
-
•
For with both parameters unknown, with inverse .
-
•
The diagonal of the inverse dominates the reciprocal of the diagonal, with equality for orthogonal parameters: nuisance parameters inflate the bound by in the two-parameter case, which is one for the normal distribution and about for the gamma distribution with shape .
Let be i.i.d. and consider the parametrisation , where the standard deviation is used.
- 1.
-
2.
By taking derivatives of the log-likelihood twice and taking expectations, verify this result directly.
-
3.
Write down the Cramer–Rao bound for an unbiased estimator for . Using Jensen’s inequality, lemma A.7, show that the sample standard deviation is not unbiased for and thus cannot be compared to directly.
Consider the gamma distribution with shape and rate . The log-likelihood of observing one value is from exercise 7.3, and the information matrix is given in section 14.5.
-
1.
Compute the first and second partial derivatives of and check the matrix.
-
2.
Assume that the shape parameter is known to be . In this case, is the only parameter. From the matrix derive the Fisher information about . Why is this a diagonal element of and not of the inverse matrix?
- 3.
-
4.
Now assume that both parameters are unknown. Show that the Cramer–Rao bound for an unbiased estimator for from a sample of size is given by
and determine the ratio of this bound to the bound for known shape parameter .
Let be i.i.d. with distribution and, independently of these, let be i.i.d. with distribution , where the common variance is known and is unknown.
-
1.
Write down the log-likelihood of all observations and show that the information matrix is .
-
2.
Using theorem 14.5, find the Cramer–Rao bound for an unbiased estimator for the difference . Show that achieves this bound.
-
3.
Assume now that is also unknown, i.e. when we have . Show that the information matrix is still diagonal and discuss the implications for the estimation of when is unknown.