Lecture 14 Multiparameter Fisher Information

In lectures 11 and 13 the parameter was a single number, but most models we actually fit, e.g. the normal distribution with unknown mean and variance, have several parameters. In this lecture we will extend the concepts of Fisher information and the Cramer–Rao bound to a vector parameter θ=(θ1,…,θk)\theta=(\theta_{1},\dots,\theta_{k}): the score will be a vector, the information will be a k×kk\times k matrix, and the bound will be an inequality of covariance matrices. The new feature here is that we will need to invert the matrix, and the inverse will give the cost of estimating one parameter in the presence of the other parameters.

14.1 The Score Vector and the Information Matrix

Setting the stage as in lecture 11, we now assume that the parameter space Θ⊆ℝk\Theta\subseteq\mathbb{R}^{k} is an open set and that the derivatives with respect to θ\theta are partial derivatives, denoted by ∂j\partial_{j} for the derivative with respect to θj\theta_{j}. The data X1,…,XnX_{1},\dots,X_{n} are now i.i.d. with density or probability weights f⁢(x;θ)f(x;\theta) and the log-likelihood function is ℓ⁢(θ)=∑i=1nℓi⁢(θ)\ell(\theta)=\sum_{i=1}^{n}\ell_{i}(\theta), where ℓi⁢(θ)=log⁡f⁢(Xi;θ)\ell_{i}(\theta)=\log f(X_{i};\theta), and we assume that the conditions of section 11.1 hold for each coordinate, allowing for differentiation under the integral sign for any pair of coordinates, and we assume that each coordinate of the score has finite variance under ℙθ\mathbb{P}_{\theta}, so that the entries of the information matrix below are finite.

Definition 14.1.

The score vector of the sample is given by the gradient of the log-likelihood function

∇ℓ⁢(θ)=(∂1ℓ⁢(θ),…,∂kℓ⁢(θ)),\nabla\ell(\theta)=\bigl{(}\partial_{1}\ell(\theta),\dots,\partial_{k}\ell(% \theta)\bigr{)},

as a column vector in ℝk\mathbb{R}^{k}.

For k=1k=1 the score vector is the score ℓ′⁢(θ)\ell^{\prime}(\theta) from definition 4.4. Setting the score vector to zero leads to the system of likelihood equations for the normal distribution considered in example 5.4. As in lecture 11, we will consider the score vector to be a random vector under ℙθ\mathbb{P}_{\theta} and we will determine the covariance of this random vector. The covariance matrix Cov(Y)\mathop{\mathrm{Cov}}\nolimits(Y) of a random vector Y=(Y1,…,Yk)Y=(Y_{1},\dots,Y_{k}) has entries Cov(Yj,Yl)\mathop{\mathrm{Cov}}\nolimits(Y_{j},Y_{l}). The diagonal of this matrix contains the variances.

Definition 14.2.

Given the conditions of section 11.1, the Fisher information matrix of the sample about θ\theta is defined as

ℐ︀n⁢(θ)=Covθ(∇ℓ⁢(θ)),i.e.ℐ︀n⁢(θ)j⁢l=Covθ(∂jℓ⁢(θ),∂lℓ⁢(θ)),\mathcal{I}_{n}(\theta)=\mathop{\mathrm{Cov}}\nolimits_{\theta}\bigl{(}\nabla% \ell(\theta)\bigr{)},\qquad\text{{i.e.}}\qquad\mathcal{I}_{n}(\theta)_{jl}=% \mathop{\mathrm{Cov}}\nolimits_{\theta}\bigl{(}\partial_{j}\ell(\theta),% \partial_{l}\ell(\theta)\bigr{)},

where this is the k×kk\times k covariance matrix of the score vector. The information matrix for a single observation is denoted by ℐ︀⁢(θ)=ℐ︀1⁢(θ)\mathcal{I}(\theta)=\mathcal{I}_{1}(\theta).

The diagonal element ℐ︀n⁢(θ)j⁢j\mathcal{I}_{n}(\theta)_{jj} is the Fisher information about θj\theta_{j} in the sense of definition 11.2 with all other coordinates fixed to their true values; the off-diagonal elements are new and measure the strength of correlations between the components of the score. Both identities from lecture 11 carry over to this situation, separately for each coordinate.

Proposition 14.3.

With the assumptions of section 11.1 in place, we have 𝔼θ⁢(∇ℓ⁢(θ))=0\mathbb{E}_{\theta}\bigl{(}\nabla\ell(\theta)\bigr{)}=0 and

ℐ︀n⁢(θ)j⁢l=𝔼θ⁢(∂jℓ⁢(θ)⁢∂lℓ⁢(θ))=−𝔼θ⁢(∂j∂lℓ⁢(θ))for ⁢j,l=1,…,k.\mathcal{I}_{n}(\theta)_{jl}=\mathbb{E}_{\theta}\bigl{(}\partial_{j}\ell(% \theta)\,\partial_{l}\ell(\theta)\bigr{)}=-\mathbb{E}_{\theta}\bigl{(}\partial% _{j}\partial_{l}\ell(\theta)\bigr{)}\qquad\text{for }j,l=1,\dots,k.

Thus, ℐ︀n⁢(θ)\mathcal{I}_{n}(\theta) is minus the expected Hessian matrix of the log-likelihood.

The proof of this result is similar to the proof of lemma 11.1 and theorem 11.3, but with partial derivatives instead of d/d⁢θ\mathrm{d}/\mathrm{d}\theta. We omit the details of the proof here, but we can mention that the ratio ∂jf/f\partial_{j}f/f has mean zero since the density in the numerator and denominator cancels. Using the quotient rule we find ∂j∂lℓi=∂j∂lf⁢(Xi;θ)/f⁢(Xi;θ)−∂jℓi⁢∂lℓi\partial_{j}\partial_{l}\ell_{i}=\partial_{j}\partial_{l}f(X_{i};\theta)/f(X_{% i};\theta)-\partial_{j}\ell_{i}\,\partial_{l}\ell_{i} where the first term has mean zero. Finally, the independence of the observations allows us to carry over the resulting identity from one observation to the sample. As for the scalar case, the second expression is the more practical one, at least for k=2k=2, because it requires only three second derivatives (one mixed derivative, occurring twice by symmetry).

14.2 Properties of the Information Matrix

The three properties of the scalar information which we used most often in lecture 11 have matrix counterparts. Recall that a symmetric matrix AA is positive semi-definite, if a⊤⁢A⁢a≥0a^{\top}Aa\geq 0 for all a∈ℝka\in\mathbb{R}^{k}, and positive definite, if a⊤⁢A⁢a>0a^{\top}Aa>0 for all a≠0a\neq 0; a positive definite matrix has a positive definite inverse.

Proposition 14.4.

Suppose the conditions from section 11.1 hold. Then the following statements hold.

  1. 1.
    ​

    The matrix ℐ︀n⁢(θ)\mathcal{I}_{n}(\theta) is symmetric and positive semi-definite. In particular we have a⊤⁢ℐ︀n⁢(θ)⁢a=Varθ(a⊤⁢∇ℓ⁢(θ))a^{\top}\mathcal{I}_{n}(\theta)a=\mathop{\mathrm{Var}}\nolimits_{\theta}\bigl{% (}a^{\top}\nabla\ell(\theta)\bigr{)} for all a∈ℝka\in\mathbb{R}^{k}.

  2. 2.
    ​

    Information is additive: ℐ︀n⁢(θ)=n⁢ℐ︀⁢(θ)\mathcal{I}_{n}(\theta)=n\,\mathcal{I}(\theta).

  3. 3.
    ​

    Let φ↦θ=h⁢(φ)\varphi\mapsto\theta=h(\varphi) be a continuously differentiable bijection between open subsets of ℝk\mathbb{R}^{k} — here φ\varphi is a parameter vector, the several-parameter counterpart of the ψ=g⁢(θ)\psi=g(\theta) in lecture 5 — and let J⁢(φ)J(\varphi) be the invertible Jacobian matrix with entries Jj⁢l=∂θj/∂φlJ_{jl}=\partial\theta_{j}/\partial\varphi_{l}. Then the information matrix for φ\varphi is given by

    ℐ︀nφ⁢(φ)=J⁢(φ)⊤⁢ℐ︀n⁢(θ)⁢J⁢(φ)for ⁢θ=h⁢(φ).\mathcal{I}^{\varphi}_{n}(\varphi)=J(\varphi)^{\top}\,\mathcal{I}_{n}(\theta)% \,J(\varphi)\qquad\text{for }\theta=h(\varphi).
Proof.

For the first statement, the symmetry of the matrix follows from the fact that Cov(Yj,Yl)=Cov(Yl,Yj)\mathop{\mathrm{Cov}}\nolimits(Y_{j},Y_{l})=\mathop{\mathrm{Cov}}\nolimits(Y_{% l},Y_{j}) for all jj and ll. For a∈ℝka\in\mathbb{R}^{k}, the variance of the linear combination a⊤⁢∇ℓ⁢(θ)a^{\top}\nabla\ell(\theta) is given by

Varθ(∑j=1kaj⁢∂jℓ⁢(θ))=∑j=1k∑l=1kaj⁢al⁢ℐ︀n⁢(θ)j⁢l=a⊤⁢ℐ︀n⁢(θ)⁢a,\mathop{\mathrm{Var}}\nolimits_{\theta}\Bigl{(}\sum_{j=1}^{k}a_{j}\,\partial_{% j}\ell(\theta)\Bigr{)}=\sum_{j=1}^{k}\sum_{l=1}^{k}a_{j}a_{l}\,\mathcal{I}_{n}% (\theta)_{jl}=a^{\top}\mathcal{I}_{n}(\theta)a,

and thus is non-negative. The second statement is proved by using the fact that the score vector of the sample is the sum of nn i.i.d. score vectors of the individual observations and that the covariance matrix of the sum of independent random vectors equals the sum of the covariance matrices. Finally, for the third statement, we note that the log-likelihood in the new parametrisation is ℓ~⁢(φ)=ℓ⁢(h⁢(φ))\tilde{\ell}(\varphi)=\ell\bigl{(}h(\varphi)\bigr{)} and thus we find

∂ℓ~∂φl⁢(φ)=∑j=1k∂jℓ⁢(θ)⁢∂θj∂φl=∑j=1kJj⁢l⁢∂jℓ⁢(θ),\frac{\partial\tilde{\ell}}{\partial\varphi_{l}}(\varphi)=\sum_{j=1}^{k}% \partial_{j}\ell(\theta)\,\frac{\partial\theta_{j}}{\partial\varphi_{l}}=\sum_% {j=1}^{k}J_{jl}\,\partial_{j}\ell(\theta),

i.e. ∇ℓ~⁢(φ)=J⊤⁢∇ℓ⁢(θ)\nabla\tilde{\ell}(\varphi)=J^{\top}\nabla\ell(\theta). Since JJ is a fixed matrix for fixed φ\varphi, the covariance matrix of J⊤⁢∇ℓ⁢(θ)J^{\top}\nabla\ell(\theta) is given by J⊤⁢ℐ︀n⁢(θ)⁢JJ^{\top}\mathcal{I}_{n}(\theta)J. This completes the proof. ∎

For k=1k=1 the third statement is a consequence of proposition 11.10: If ψ=g⁢(θ)\psi=g(\theta) we have h=g−1h=g^{-1} and J=1/g′⁢(θ)J=1/g^{\prime}(\theta). Positive definiteness is violated in cases where the log-likelihood does not respond to changes of θ\theta in some direction, for example when the model is not identifiable in the sense of definition 4.10. From now on we assume that ℐ︀n⁢(θ)\mathcal{I}_{n}(\theta) is positive definite so that the inverse exists.

14.3 The Cramer–Rao Bound for Vector Parameters

Theorem 13.2 bounds the variance of an unbiased estimator for a scalar parameter from below. For a vector parameter the natural object to bound is the covariance matrix, and the natural reading of “at least” between symmetric matrices is that the difference is positive semi-definite.

Theorem 14.5.

Suppose we have the conditions of section 11.1 and a positive definite matrix ℐ︀n⁢(θ)\mathcal{I}_{n}(\theta). Let T=(T1,…,Tk)T=(T_{1},\dots,T_{k}) be an unbiased estimator for θ\theta, so that 𝔼θ⁢(Tj)=θj\mathbb{E}_{\theta}(T_{j})=\theta_{j} for all jj and all θ\theta, where the variances are finite. Then the matrix

Covθ(T)−ℐ︀n⁢(θ)−1\mathop{\mathrm{Cov}}\nolimits_{\theta}(T)-\mathcal{I}_{n}(\theta)^{-1}

is positive semi-definite for all θ∈Θ\theta\in\Theta. In particular we have

Varθ(Tj)≥(ℐ︀n⁢(θ)−1)j⁢jandVarθ(a⊤⁢T)≥a⊤⁢ℐ︀n⁢(θ)−1⁢a\mathop{\mathrm{Var}}\nolimits_{\theta}(T_{j})\geq\bigl{(}\mathcal{I}_{n}(% \theta)^{-1}\bigr{)}_{jj}\qquad\text{and}\qquad\mathop{\mathrm{Var}}\nolimits_% {\theta}(a^{\top}T)\geq a^{\top}\mathcal{I}_{n}(\theta)^{-1}a

for j=1,…,kj=1,\dots,k and for all a∈ℝka\in\mathbb{R}^{k}.

Proof.

We only sketch the argument, which is the proof of theorem 13.2 for linear combinations. Differentiating 𝔼θ⁢(Tj)=θj\mathbb{E}_{\theta}(T_{j})=\theta_{j} under the integral sign, as in that proof, gives Covθ(Tj,∂lℓ⁢(θ))=1\mathop{\mathrm{Cov}}\nolimits_{\theta}\bigl{(}T_{j},\partial_{l}\ell(\theta)% \bigr{)}=1 for l=jl=j and 0 otherwise, and thus by bilinearity Covθ(a⊤⁢T,b⊤⁢∇ℓ⁢(θ))=a⊤⁢b\mathop{\mathrm{Cov}}\nolimits_{\theta}\bigl{(}a^{\top}T,b^{\top}\nabla\ell(% \theta)\bigr{)}=a^{\top}b for all a,b∈ℝka,b\in\mathbb{R}^{k}. Using the Cauchy–Schwarz inequality, lemma A.8, and the first statement of proposition 14.4, we find

(a⊤⁢b)2≤(a⊤⁢Covθ(T)⁢a)⁢(b⊤⁢ℐ︀n⁢(θ)⁢b).(a^{\top}b)^{2}\leq\bigl{(}a^{\top}\mathop{\mathrm{Cov}}\nolimits_{\theta}(T)% \,a\bigr{)}\bigl{(}b^{\top}\mathcal{I}_{n}(\theta)\,b\bigr{)}.

Choosing b=ℐ︀n⁢(θ)−1⁢ab=\mathcal{I}_{n}(\theta)^{-1}a, both a⊤⁢ba^{\top}b and b⊤⁢ℐ︀n⁢(θ)⁢bb^{\top}\mathcal{I}_{n}(\theta)b equal a⊤⁢ℐ︀n⁢(θ)−1⁢aa^{\top}\mathcal{I}_{n}(\theta)^{-1}a, which is positive for a≠0a\neq 0, and dividing by this value gives a⊤⁢Covθ(T)⁢a≥a⊤⁢ℐ︀n⁢(θ)−1⁢aa^{\top}\mathop{\mathrm{Cov}}\nolimits_{\theta}(T)\,a\geq a^{\top}\mathcal{I}_% {n}(\theta)^{-1}a. This is the claim, since a⊤⁢Covθ(T)⁢a=Varθ(a⊤⁢T)a^{\top}\mathop{\mathrm{Cov}}\nolimits_{\theta}(T)\,a=\mathop{\mathrm{Var}}% \nolimits_{\theta}(a^{\top}T), and for the diagonal case we have a=eja=e_{j}, with the jjth unit vector. This completes the sketch of the proof. ∎

The bound is almost always applied in its diagonal form: The variance of any unbiased estimator for θj\theta_{j}, when the other components are unknown, is at least the jjth diagonal entry of the inverse information matrix. Note that this is the diagonal of the inverse, not the reciprocal of the diagonal; we will see the difference in section 14.5. An estimator which achieves the bound for θj\theta_{j} is said to be efficient for θj\theta_{j}, as in lecture 13. The bound for a function g⁢(θ)g(\theta) with gradient ∇g\nabla g is ∇g⁢(θ)⊤⁢ℐ︀n⁢(θ)−1⁢∇g⁢(θ)\nabla g(\theta)^{\top}\mathcal{I}_{n}(\theta)^{-1}\nabla g(\theta), which includes the statement of theorem 13.3 for k=1k=1, as well as the linear case g⁢(θ)=a⊤⁢θg(\theta)=a^{\top}\theta, which the theorem already covers. The result is also used in exercise 14.3.

14.4 The Normal Distribution

The only example of a distribution which all students should be able to reproduce without notes is the normal distribution with unknown parameters. We will consider this example in detail.

Example 14.6.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. N⁢(μ,σ2)N(\mu,\sigma^{2}) where θ=(μ,σ2)\theta=(\mu,\sigma^{2}), both unknown. From example 5.4 we know that the log-likelihood function is

ℓ⁢(μ,σ2)=−n2⁢log⁡(2⁢π)−n2⁢log⁡σ2−12⁢σ2⁢∑i=1n(Xi−μ)2,\ell(\mu,\sigma^{2})=-\frac{n}{2}\log(2\pi)-\frac{n}{2}\log\sigma^{2}-\frac{1}% {2\sigma^{2}}\sum_{i=1}^{n}(X_{i}-\mu)^{2},

and the first partial derivatives are given by

∂ℓ∂μ=1σ2⁢∑i=1n(Xi−μ),∂ℓ∂σ2=−n2⁢σ2+12⁢σ4⁢∑i=1n(Xi−μ)2.\frac{\partial\ell}{\partial\mu}=\frac{1}{\sigma^{2}}\sum_{i=1}^{n}(X_{i}-\mu)% ,\qquad\frac{\partial\ell}{\partial\sigma^{2}}=-\frac{n}{2\sigma^{2}}+\frac{1}% {2\sigma^{4}}\sum_{i=1}^{n}(X_{i}-\mu)^{2}.

Taking second derivatives, where we treat σ2\sigma^{2} as a single variable, we find the three second derivatives as

∂2ℓ∂μ2\displaystyle\frac{\partial^{2}\ell}{\partial\mu^{2}}
=−nσ2,∂2ℓ∂μ⁢∂σ2=−1σ4⁢∑i=1n(Xi−μ),\displaystyle=-\frac{n}{\sigma^{2}},\qquad\frac{\partial^{2}\ell}{\partial\mu% \,\partial\sigma^{2}}=-\frac{1}{\sigma^{4}}\sum_{i=1}^{n}(X_{i}-\mu),
∂2ℓ∂(σ2)2\displaystyle\frac{\partial^{2}\ell}{\partial(\sigma^{2})^{2}}
=n2⁢σ4−1σ6⁢∑i=1n(Xi−μ)2.\displaystyle=\frac{n}{2\sigma^{4}}-\frac{1}{\sigma^{6}}\sum_{i=1}^{n}(X_{i}-% \mu)^{2}.

To compute the expectations, we use the fact that 𝔼θ⁢(Xi−μ)=0\mathbb{E}_{\theta}(X_{i}-\mu)=0 and 𝔼θ⁢((Xi−μ)2)=σ2\mathbb{E}_{\theta}\bigl{(}(X_{i}-\mu)^{2}\bigr{)}=\sigma^{2}. The mixed derivative has expectation zero. The last term has expectation n/(2⁢σ4)−n⁢σ2/σ6=−n/(2⁢σ4)n/(2\sigma^{4})-n\sigma^{2}/\sigma^{6}=-n/(2\sigma^{4}). Thus, by proposition 14.3, we find the information matrix and its inverse to be

equation (14.1) (14.1)
ℐ︀n⁢(μ,σ2)=(nσ200n2⁢σ4)andℐ︀n⁢(μ,σ2)−1=(σ2n002⁢σ4n).\mathcal{I}_{n}(\mu,\sigma^{2})=\begin{pmatrix}\dfrac{n}{\sigma^{2}}&0\\[9.429% 14pt] 0&\dfrac{n}{2\sigma^{4}}\end{pmatrix}\qquad\text{and}\qquad\mathcal{I}_{n}(\mu% ,\sigma^{2})^{-1}=\begin{pmatrix}\dfrac{\sigma^{2}}{n}&0\\[9.42914pt] 0&\dfrac{2\sigma^{4}}{n}\end{pmatrix}.

We should commit these matrices to memory. The upper left entry is the information n/σ2n/\sigma^{2} about the mean, where the variance is known from example 11.7. The sample mean, with variance σ2/n\sigma^{2}/n, achieves the bound for μ\mu. For the variance the bound is 2⁢σ4/n2\sigma^{4}/n. The UMVUE S2S^{2} from example 10.7 has variance 2⁢σ4/(n−1)2\sigma^{4}/(n-1), as shown in the computation of section 2.3, and thus does not achieve the bound. Since S2S^{2} is the UMVUE, no unbiased estimator for σ2\sigma^{2} achieves the bound when μ\mu is unknown. If we had parametrised the model by the standard deviation instead, i.e. if we had φ=(μ,σ)\varphi=(\mu,\sigma), then the third statement of proposition 14.4, with J=diag⁢(1,2⁢σ)J=\mathrm{diag}(1,2\sigma), gives ℐ︀nφ⁢(μ,σ)=diag⁢(n/σ2, 2⁢n/σ2)\mathcal{I}^{\varphi}_{n}(\mu,\sigma)=\mathrm{diag}(n/\sigma^{2},\,2n/\sigma^{% 2}). See exercise 14.1 for details.

The information matrix for the normal distribution is diagonal. The following section explains the significance of this fact and what happens when this is not the case.

14.5 Nuisance Parameters

Assume that we are interested in θ1\theta_{1} only, and that the remaining components θ2,…,θk\theta_{2},\dots,\theta_{k} are nuisance parameters, i.e. unknown but not of interest. If the values of the nuisance parameters were known, the Cramer–Rao bound for θ1\theta_{1} would be 1/ℐ︀n⁢(θ)111/\mathcal{I}_{n}(\theta)_{11}. Since the values are unknown, the bound is (ℐ︀n⁢(θ)−1)11\bigl{(}\mathcal{I}_{n}(\theta)^{-1}\bigr{)}_{11} by theorem 14.5. For k=2k=2 we can compare the two bounds explicitly. For this we write

ℐ︀n⁢(θ)=(abbc)and findℐ︀n⁢(θ)−1=1a⁢c−b2⁢(c−b−ba),\mathcal{I}_{n}(\theta)=\begin{pmatrix}a&b\\ b&c\end{pmatrix}\qquad\text{and find}\qquad\mathcal{I}_{n}(\theta)^{-1}=\frac{% 1}{ac-b^{2}}\begin{pmatrix}c&-b\\ -b&a\end{pmatrix},

where a⁢c−b2>0ac-b^{2}>0 since the matrix is positive definite, and

equation (14.2) (14.2)
(ℐ︀n⁢(θ)−1)11=ca⁢c−b2=1a⋅11−ρ2,where ⁢ρ=ba⁢c\bigl{(}\mathcal{I}_{n}(\theta)^{-1}\bigr{)}_{11}=\frac{c}{ac-b^{2}}=\frac{1}{% a}\cdot\frac{1}{1-\rho^{2}},\qquad\text{where }\rho=\frac{b}{\sqrt{ac}}

is the correlation between the two components of the score vector. Since 0≤ρ2<10\leq\rho^{2}<1, the bound where the nuisance parameter is unknown is 1/(1−ρ2)1/(1-\rho^{2}) times larger than the bound where the nuisance parameter is known, and both bounds are equal if and only if b=0b=0. Two parameters with ℐ︀n⁢(θ)12=0\mathcal{I}_{n}(\theta)_{12}=0 are said to be orthogonal; in this case, not knowing θ2\theta_{2} does not incur any additional cost for the bound when estimating θ1\theta_{1}. The same inequality (ℐ︀n⁢(θ)−1)j⁢j≥1/ℐ︀n⁢(θ)j⁢j\bigl{(}\mathcal{I}_{n}(\theta)^{-1}\bigr{)}_{jj}\geq 1/\mathcal{I}_{n}(\theta% )_{jj} holds for all jj and all kk: not knowing a parameter can never be beneficial.

The normal distribution is the orthogonal case: in example 14.6 the mixed derivative has mean zero, and the bound for μ\mu is σ2/n\sigma^{2}/n whether or not σ2\sigma^{2} is known. This is a special feature of the normal distribution, not the rule. For the gamma distribution with shape α>0\alpha>0 and rate β>0\beta>0 both unknown, as in exercise 7.3, all three second derivatives of the log-likelihood ℓ1⁢(α,β)\ell_{1}(\alpha,\beta) of one observation are non-random, and the information matrix of one observation is

ℐ︀⁢(α,β)=(ψ′⁢(α)−1β−1βαβ2),\mathcal{I}(\alpha,\beta)=\begin{pmatrix}\psi^{\prime}(\alpha)&-\dfrac{1}{% \beta}\\[9.42914pt] -\dfrac{1}{\beta}&\dfrac{\alpha}{\beta^{2}}\end{pmatrix},

where ψ′\psi^{\prime}, the second derivative of log⁡Γ\log\Gamma, is the trigamma function; exercise 14.2 derives this matrix. The off-diagonal entry is non-zero and by  (14.2) we have ρ2=1/(α⁢ψ′⁢(α))\rho^{2}=1/\bigl{(}\alpha\psi^{\prime}(\alpha)\bigr{)}; for α=2\alpha=2, with ψ′⁢(2)=π2/6−1≈0.645\psi^{\prime}(2)=\pi^{2}/6-1\approx 0.645, this gives 1/(1−ρ2)≈4.451/(1-\rho^{2})\approx 4.45, so that the bound with the other parameter unknown is about four and a half times the bound with the other parameter known.

The inverse information matrix is needed when we estimate several parameters simultaneously. In lecture 17 we will see that, for large sample size, the maximum likelihood estimator is approximately normally distributed with covariance matrix ℐ︀n⁢(θ0)−1\mathcal{I}_{n}(\theta_{0})^{-1}. In lecture 18 we will see how we can numerically find the standard errors of the estimator by inverting the observed Hessian matrix, and in lecture 23 we will see that the number of parameters kk can be seen as the number of degrees of freedom in this system.

Summary.
  • •
    ​

    For a vector parameter the score is the gradient ∇ℓ⁢(θ)\nabla\ell(\theta), and the Fisher information matrix ℐ︀n⁢(θ)\mathcal{I}_{n}(\theta) is its covariance matrix, equal to minus the expected Hessian of the log-likelihood.

  • •
    ​

    The information matrix is symmetric and positive semi-definite, adds over independent observations, ℐ︀n=n⁢ℐ︀\mathcal{I}_{n}=n\mathcal{I}, and transforms as J⊤⁢ℐ︀n⁢JJ^{\top}\mathcal{I}_{n}J under a change of parametrisation with Jacobian JJ.

  • •
    ​

    Cramer–Rao: for an unbiased estimator TT for θ\theta the matrix Covθ(T)−ℐ︀n⁢(θ)−1\mathop{\mathrm{Cov}}\nolimits_{\theta}(T)-\mathcal{I}_{n}(\theta)^{-1} is positive semi-definite; in particular Varθ(Tj)≥(ℐ︀n⁢(θ)−1)j⁢j\mathop{\mathrm{Var}}\nolimits_{\theta}(T_{j})\geq\bigl{(}\mathcal{I}_{n}(% \theta)^{-1}\bigr{)}_{jj}.

  • •
    ​

    For N⁢(μ,σ2)N(\mu,\sigma^{2}) with both parameters unknown, ℐ︀n=diag⁢(n/σ2,n/(2⁢σ4))\mathcal{I}_{n}=\mathrm{diag}\bigl{(}n/\sigma^{2},\,n/(2\sigma^{4})\bigr{)} with inverse diag⁢(σ2/n, 2⁢σ4/n)\mathrm{diag}(\sigma^{2}/n,\,2\sigma^{4}/n).

  • •
    ​

    The diagonal of the inverse dominates the reciprocal of the diagonal, with equality for orthogonal parameters: nuisance parameters inflate the bound by 1/(1−ρ2)1/(1-\rho^{2}) in the two-parameter case, which is one for the normal distribution and about 4.454.45 for the gamma distribution with shape 22.

Exercise 14.1.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. N⁢(μ,σ2)N(\mu,\sigma^{2}) and consider the parametrisation φ=(μ,σ)\varphi=(\mu,\sigma), where the standard deviation is used.

  1. 1.
    ​

    Compute the Jacobian matrix JJ of (μ,σ)↦(μ,σ2)(\mu,\sigma)\mapsto(\mu,\sigma^{2}) and use proposition 14.4 and (14.1) to show that

    ℐ︀nφ⁢(μ,σ)=diag⁢(n/σ2, 2⁢n/σ2).\mathcal{I}^{\varphi}_{n}(\mu,\sigma)=\mathrm{diag}(n/\sigma^{2},\,2n/\sigma^{% 2}).
  2. 2.
    ​

    By taking derivatives of the log-likelihood ℓ⁢(μ,σ)\ell(\mu,\sigma) twice and taking expectations, verify this result directly.

  3. 3.
    ​

    Write down the Cramer–Rao bound for an unbiased estimator for σ\sigma. Using Jensen’s inequality, lemma A.7, show that the sample standard deviation S=S2S=\sqrt{S^{2}} is not unbiased for σ\sigma and thus cannot be compared to Var(S)\mathop{\mathrm{Var}}\nolimits(S) directly.

Exercise 14.2.

Consider the gamma distribution with shape α>0\alpha>0 and rate β>0\beta>0. The log-likelihood of observing one value is ℓ1⁢(α,β)=α⁢log⁡β−log⁡Γ⁢(α)+(α−1)⁢log⁡X−β⁢X\ell_{1}(\alpha,\beta)=\alpha\log\beta-\log\Gamma(\alpha)+(\alpha-1)\log X-\beta X from exercise 7.3, and the information matrix ℐ︀⁢(α,β)\mathcal{I}(\alpha,\beta) is given in section 14.5.

  1. 1.
    ​

    Compute the first and second partial derivatives of ℓ1⁢(α,β)\ell_{1}(\alpha,\beta) and check the matrix.

  2. 2.
    ​

    Assume that the shape parameter is known to be α=2\alpha=2. In this case, β\beta is the only parameter. From the matrix derive the Fisher information about β\beta. Why is this a diagonal element of ℐ︀⁢(α,β)\mathcal{I}(\alpha,\beta) and not of the inverse matrix?

  3. 3.
    ​

    Exercise 11.2 considers parametrisation of the same model in terms of the scale parameter θ=1/β\theta=1/\beta. Use proposition 11.10 to derive the value ℐ︀⁢(θ)=2/θ2\mathcal{I}(\theta)=2/\theta^{2} found in that exercise.

  4. 4.
    ​

    Now assume that both parameters are unknown. Show that the Cramer–Rao bound for an unbiased estimator for β\beta from a sample of size nn is given by

    β2⁢ψ′⁢(α)n⁢(α⁢ψ′⁢(α)−1),\frac{\beta^{2}\psi^{\prime}(\alpha)}{n\bigl{(}\alpha\psi^{\prime}(\alpha)-1% \bigr{)}},

    and determine the ratio of this bound to the bound for known shape parameter α=2\alpha=2.

Exercise 14.3.

Let X1,…,XmX_{1},\dots,X_{m} be i.i.d. with distribution N⁢(μ1,σ2)N(\mu_{1},\sigma^{2}) and, independently of these, let Y1,…,YnY_{1},\dots,Y_{n} be i.i.d. with distribution N⁢(μ2,σ2)N(\mu_{2},\sigma^{2}), where the common variance σ2\sigma^{2} is known and θ=(μ1,μ2)\theta=(\mu_{1},\mu_{2}) is unknown.

  1. 1.
    ​

    Write down the log-likelihood of all m+nm+n observations and show that the information matrix is ℐ︀⁢(μ1,μ2)=diag⁢(m/σ2,n/σ2)\mathcal{I}(\mu_{1},\mu_{2})=\mathrm{diag}(m/\sigma^{2},\,n/\sigma^{2}).

  2. 2.
    ​

    Using theorem 14.5, find the Cramer–Rao bound for an unbiased estimator for the difference μ1−μ2\mu_{1}-\mu_{2}. Show that X¯−Y¯\bar{X}-\bar{Y} achieves this bound.

  3. 3.
    ​

    Assume now that σ2\sigma^{2} is also unknown, i.e. when we have θ=(μ1,μ2,σ2)\theta=(\mu_{1},\mu_{2},\sigma^{2}). Show that the 3×33\times 3 information matrix is still diagonal and discuss the implications for the estimation of μ1−μ2\mu_{1}-\mu_{2} when σ2\sigma^{2} is unknown.