Lecture 11 Fisher Information

Since the data are random, the score function from lecture 4 is a random variable, and the Fisher information, given by the variance of the score, measures how much information the sample contains about the parameter. In lecture 13 we will use the Fisher information to obtain a lower bound for the variance of any unbiased estimator, and in lecture 17 we will see that the Fisher information equals the asymptotic variance of the maximum likelihood estimator. In this lecture we will define the Fisher information, prove two identities which allow us to compute the Fisher information, and we will compute the Fisher information for the standard families of models.

11.1 Setting and Assumptions

Throughout this lecture, X1,…,XnX_{1},\dots,X_{n} are i.i.d. with density or probability weights f⁢(x;θ)f(x;\theta), where the parameter space Θ⊆ℝ\Theta\subseteq\mathbb{R} is an open interval. As in lecture 4, the log-likelihood is the sum ℓ⁢(θ)=∑i=1nℓi⁢(θ)\ell(\theta)=\sum_{i=1}^{n}\ell_{i}(\theta) of the terms ℓi⁢(θ)=log⁡f⁢(Xi;θ)\ell_{i}(\theta)=\log f(X_{i};\theta) and the score of definition 4.4 is its derivative ℓ′⁢(θ)=∑iℓi′⁢(θ)\ell^{\prime}(\theta)=\sum_{i}\ell_{i}^{\prime}(\theta). In this section we study the distribution of the score under ℙθ\mathbb{P}_{\theta}, for a fixed value of θ\theta.

The results of this lecture are obtained by differentiating, with respect to θ\theta, the identity ∫f⁢(x;θ)⁢dx=1\int f(x;\theta)\,\mathrm{d}x=1, and we need this differentiation to be legitimate. We assume the following throughout. The support {x∣f⁢(x;θ)>0}\{x\mid f(x;\theta)>0\} of the distribution does not depend on θ\theta. For every xx the function θ↦f⁢(x;θ)\theta\mapsto f(x;\theta) is twice continuously differentiable on Θ\Theta. Finally, the integral of f⁢(x;θ)f(x;\theta) over the support may be differentiated under the integral sign twice, so that

equation (11.1) (11.1)
dkd⁢θk⁢∫f⁢(x;θ)⁢dx=∫∂k∂θk⁢f⁢(x;θ)⁢dxfor ⁢k=1,2,\frac{\mathrm{d}^{k}}{\mathrm{d}\theta^{k}}\int f(x;\theta)\,\mathrm{d}x=\int% \frac{\partial^{k}}{\partial\theta^{k}}f(x;\theta)\,\mathrm{d}x\qquad\text{for% }k=1,2,

with sums in place of integrals for a discrete model. These three assumptions are referred to as the conditions of this section. Most standard families of models from tables 1.1 and 1.2 satisfy these conditions, except for the uniform distribution (in section 11.7). The exponential families in the sense of definition 7.1 with twice continuously differentiable natural parameter η⁢(θ)\eta(\theta) also satisfy these conditions (since the exchange is the step in the proof of theorem 7.7). These conditions are included in the full list of regularity conditions in definition 16.1 of lecture 16.

11.2 The Score Has Mean Zero

The first identity is that, if θ\theta is the correct parameter value, the data do not systematically deviate from this value: the score at the true parameter value is zero on average.

Lemma 11.1.

Let the situation be as described in section 11.1. Then

𝔼θ⁢(ℓi′⁢(θ))=0for ⁢i=1,…,n,and thus𝔼θ⁢(ℓ′⁢(θ))=0\mathbb{E}_{\theta}\bigl{(}\ell_{i}^{\prime}(\theta)\bigr{)}=0\qquad\text{for % }i=1,\dots,n,\qquad\text{and thus}\qquad\mathbb{E}_{\theta}\bigl{(}\ell^{% \prime}(\theta)\bigr{)}=0

for all θ∈Θ\theta\in\Theta.

Proof.

Using the chain rule, we find the derivative of ℓi⁢(θ)=log⁡f⁢(Xi;θ)\ell_{i}(\theta)=\log f(X_{i};\theta) as

equation (11.2) (11.2)
ℓi′⁢(θ)=∂∂θ⁢f⁢(Xi;θ)f⁢(Xi;θ).\ell_{i}^{\prime}(\theta)=\frac{\frac{\partial}{\partial\theta}f(X_{i};\theta)% }{f(X_{i};\theta)}.

Since the denominator is positive (because XiX_{i} is in the support), we can cancel the density with the denominator to find

𝔼θ⁢(ℓi′⁢(θ))\displaystyle\mathbb{E}_{\theta}\bigl{(}\ell_{i}^{\prime}(\theta)\bigr{)}
=∫∂∂θ⁢f⁢(x;θ)f⁢(x;θ)⁢f⁢(x;θ)⁢dx=∫∂∂θ⁢f⁢(x;θ)⁢dx\displaystyle=\int\frac{\frac{\partial}{\partial\theta}f(x;\theta)}{f(x;\theta% )}\,f(x;\theta)\,\mathrm{d}x=\int\frac{\partial}{\partial\theta}f(x;\theta)\,% \mathrm{d}x
=dd⁢θ⁢∫f⁢(x;θ)⁢dx=dd⁢θ⁢ 1=0,\displaystyle=\frac{\mathrm{d}}{\mathrm{d}\theta}\int f(x;\theta)\,\mathrm{d}x% =\frac{\mathrm{d}}{\mathrm{d}\theta}\,1=0,

where the integral is over the support of the distribution, which does not depend on θ\theta, and the third equality comes from the case k=1k=1 of (11.1). Summing over ii gives the statement for the whole sample. This completes the proof. ∎

If the parameter value is wrong, the score is not mean zero: for the exponential model from example 4.5, the score has expectation n/θ−n/θ0n/\theta-n/\theta_{0} if the true parameter value is θ0\theta_{0}, positive for θ<θ0\theta<\theta_{0} and negative for θ>θ0\theta>\theta_{0}. The score on average points towards the correct parameter value from both sides. This is the mechanism behind the consistency proof of lecture 16.

11.3 Fisher Information

The next question we could ask is about the fluctuations of the score around zero. If the score is typically large in magnitude, this indicates that the data allow us to distinguish well between nearby parameter values. In contrast, if the score is typically close to zero, this indicates that the log-likelihood is nearly flat at the truth. The variance of the score is a measure for how informative the sample is.

Definition 11.2.

Under the conditions of section 11.1, the Fisher information of the sample X1,…,XnX_{1},\dots,X_{n} about θ\theta is given by

ℐ︀n⁢(θ)=Varθ(ℓ′⁢(θ)),\mathcal{I}_{n}(\theta)=\mathop{\mathrm{Var}}\nolimits_{\theta}\bigl{(}\ell^{% \prime}(\theta)\bigr{)},

where the variance of the score is taken at θ\theta. The Fisher information of a single observation is denoted by ℐ︀⁢(θ)=ℐ︀1⁢(θ)\mathcal{I}(\theta)=\mathcal{I}_{1}(\theta) and is given by the variance of ℓ1′⁢(θ)\ell_{1}^{\prime}(\theta).

The variance of the score is often awkward to compute directly, and the second expression of the following theorem is nearly always the quickest route to the information.

Theorem 11.3.

Suppose we are in the situation of section 11.1. Then

ℐ︀n⁢(θ)=𝔼θ⁢(ℓ′⁢(θ)2)=−𝔼θ⁢(ℓ′′⁢(θ))\mathcal{I}_{n}(\theta)=\mathbb{E}_{\theta}\bigl{(}\ell^{\prime}(\theta)^{2}% \bigr{)}=-\mathbb{E}_{\theta}\bigl{(}\ell^{\prime\prime}(\theta)\bigr{)}

for all θ∈Θ\theta\in\Theta.

Proof.

The first equality follows from the fact that ℓ′⁢(θ)\ell^{\prime}(\theta) has mean zero by lemma 11.1, and thus the variance equals the second moment. The second equality is obtained by taking the derivative of the ratio (11.2) using the quotient rule:

ℓi′′⁢(θ)=∂2∂θ2⁢f⁢(Xi;θ)f⁢(Xi;θ)−(∂∂θ⁢f⁢(Xi;θ)f⁢(Xi;θ))2=∂2∂θ2⁢f⁢(Xi;θ)f⁢(Xi;θ)−ℓi′⁢(θ)2.\ell_{i}^{\prime\prime}(\theta)=\frac{\frac{\partial^{2}}{\partial\theta^{2}}f% (X_{i};\theta)}{f(X_{i};\theta)}-\Bigl{(}\frac{\frac{\partial}{\partial\theta}% f(X_{i};\theta)}{f(X_{i};\theta)}\Bigr{)}^{2}=\frac{\frac{\partial^{2}}{% \partial\theta^{2}}f(X_{i};\theta)}{f(X_{i};\theta)}-\ell_{i}^{\prime}(\theta)% ^{2}.

Similar to the proof of lemma 11.1, the expectation of the first term cancels out in this case, using the case k=2k=2 of (11.1):

𝔼θ⁢(∂2∂θ2⁢f⁢(Xi;θ)f⁢(Xi;θ))=∫∂2∂θ2⁢f⁢(x;θ)⁢dx=d2d⁢θ2⁢∫f⁢(x;θ)⁢dx=0.\mathbb{E}_{\theta}\Bigl{(}\frac{\frac{\partial^{2}}{\partial\theta^{2}}f(X_{i% };\theta)}{f(X_{i};\theta)}\Bigr{)}=\int\frac{\partial^{2}}{\partial\theta^{2}% }f(x;\theta)\,\mathrm{d}x=\frac{\mathrm{d}^{2}}{\mathrm{d}\theta^{2}}\int f(x;% \theta)\,\mathrm{d}x=0.

Thus we have 𝔼θ⁢(ℓi′′⁢(θ))=−𝔼θ⁢(ℓi′⁢(θ)2)\mathbb{E}_{\theta}\bigl{(}\ell_{i}^{\prime\prime}(\theta)\bigr{)}=-\mathbb{E}% _{\theta}\bigl{(}\ell_{i}^{\prime}(\theta)^{2}\bigr{)} for all ii. The terms ℓ1′⁢(θ),…,ℓn′⁢(θ)\ell_{1}^{\prime}(\theta),\dots,\ell_{n}^{\prime}(\theta) are independent of each other, since they are functions of different observations, and since they have mean zero, the variance of the sum is the sum of the variances:

Varθ(ℓ′⁢(θ))\displaystyle\mathop{\mathrm{Var}}\nolimits_{\theta}\bigl{(}\ell^{\prime}(% \theta)\bigr{)}
=∑i=1nVarθ(ℓi′⁢(θ))=∑i=1n𝔼θ⁢(ℓi′⁢(θ)2)\displaystyle=\sum_{i=1}^{n}\mathop{\mathrm{Var}}\nolimits_{\theta}\bigl{(}% \ell_{i}^{\prime}(\theta)\bigr{)}=\sum_{i=1}^{n}\mathbb{E}_{\theta}\bigl{(}% \ell_{i}^{\prime}(\theta)^{2}\bigr{)}
=−∑i=1n𝔼θ⁢(ℓi′′⁢(θ))=−𝔼θ⁢(ℓ′′⁢(θ)),\displaystyle=-\sum_{i=1}^{n}\mathbb{E}_{\theta}\bigl{(}\ell_{i}^{\prime\prime% }(\theta)\bigr{)}=-\mathbb{E}_{\theta}\bigl{(}\ell^{\prime\prime}(\theta)\bigr% {)},

where we used ℓ′′⁢(θ)=∑iℓi′′⁢(θ)\ell^{\prime\prime}(\theta)=\sum_{i}\ell_{i}^{\prime\prime}(\theta) in the last step. This completes the proof. ∎

The second expression gives the name of the quantity: it is the expected curvature of the log-likelihood at the true parameter value. Large information means that the log-likelihood has a sharp maximum, well-defined by the data. Small information means the log-likelihood is flat. In lecture 17 this idea will be turned into a theorem: the approximate variance of the maximum likelihood estimator is 1/ℐ︀n⁢(θ)1/\mathcal{I}_{n}(\theta) for large sample size, and the numerical value for the curvature −ℓ′′⁢(θ^)-\ell^{\prime\prime}(\hat{\theta}) at the maximum likelihood estimate can be obtained by a numerical optimiser (lecture 18).

The proof also shows how the information of the sample relates to the information of one observation.

Proposition 11.4.

With the same assumptions as in section 11.1 we have

ℐ︀n⁢(θ)=n⁢ℐ︀⁢(θ)\mathcal{I}_{n}(\theta)=n\,\mathcal{I}(\theta)

for all θ∈Θ\theta\in\Theta.

Proof.

The score is the sum of independent terms and thus the variance can be found as the sum ∑iVarθ(ℓi′⁢(θ))\sum_{i}\mathop{\mathrm{Var}}\nolimits_{\theta}\bigl{(}\ell_{i}^{\prime}(% \theta)\bigr{)}, as in the proof of theorem 11.3, and since the observations are identically distributed, each term in the sum equals

Varθ(ℓ1′⁢(θ))=ℐ︀⁢(θ).\mathop{\mathrm{Var}}\nolimits_{\theta}\bigl{(}\ell_{1}^{\prime}(\theta)\bigr{% )}=\mathcal{I}(\theta).

This completes the proof. ∎

11.4 Examples

We now compute the information for the standard families of tables 1.1 and 1.2. By proposition 11.4 it suffices to work with a single observation X=X1X=X_{1}.

Example 11.5.

For the Bernoulli distribution with success probability θ∈(0,1)\theta\in(0,1) we have ℓ1⁢(θ)=X⁢log⁡θ+(1−X)⁢log⁡(1−θ)\ell_{1}(\theta)=X\log\theta+(1-X)\log(1-\theta) and thus

ℓ1′⁢(θ)=Xθ−1−X1−θ,ℓ1′′⁢(θ)=−Xθ2−1−X(1−θ)2.\ell_{1}^{\prime}(\theta)=\frac{X}{\theta}-\frac{1-X}{1-\theta},\qquad\ell_{1}% ^{\prime\prime}(\theta)=-\frac{X}{\theta^{2}}-\frac{1-X}{(1-\theta)^{2}}.

Using 𝔼θ⁢(X)=θ\mathbb{E}_{\theta}(X)=\theta we find

ℐ︀⁢(θ)=−𝔼θ⁢(ℓ1′′⁢(θ))=θθ2+1−θ(1−θ)2=1θ⁢(1−θ).\mathcal{I}(\theta)=-\mathbb{E}_{\theta}\bigl{(}\ell_{1}^{\prime\prime}(\theta% )\bigr{)}=\frac{\theta}{\theta^{2}}+\frac{1-\theta}{(1-\theta)^{2}}=\frac{1}{% \theta(1-\theta)}.

Thus, the information of the sample is ℐ︀n⁢(θ)=n/(θ⁢(1−θ))\mathcal{I}_{n}(\theta)=n/\bigl{(}\theta(1-\theta)\bigr{)}. Exercise 11.1 checks these identities for this model directly.

Example 11.6.

For the Poisson distribution with mean θ>0\theta>0 we have ℓ1⁢(θ)=X⁢log⁡θ−θ−log⁡X!\ell_{1}(\theta)=X\log\theta-\theta-\log X! and thus

ℓ1′⁢(θ)=Xθ−1,ℓ1′′⁢(θ)=−Xθ2,ℐ︀⁢(θ)=𝔼θ⁢(X)θ2=1θ.\ell_{1}^{\prime}(\theta)=\frac{X}{\theta}-1,\qquad\ell_{1}^{\prime\prime}(% \theta)=-\frac{X}{\theta^{2}},\qquad\mathcal{I}(\theta)=\frac{\mathbb{E}_{% \theta}(X)}{\theta^{2}}=\frac{1}{\theta}.
Example 11.7.

In the case of the normal distribution with unknown mean μ\mu and known variance σ2\sigma^{2}, the parameter is μ\mu and ℓ1⁢(μ)=−12⁢log⁡(2⁢π⁢σ2)−(X−μ)2/(2⁢σ2)\ell_{1}(\mu)=-\tfrac{1}{2}\log(2\pi\sigma^{2})-(X-\mu)^{2}/(2\sigma^{2}). Thus we have

ℓ1′⁢(μ)=X−μσ2,ℓ1′′⁢(μ)=−1σ2,ℐ︀⁢(μ)=1σ2.\ell_{1}^{\prime}(\mu)=\frac{X-\mu}{\sigma^{2}},\qquad\ell_{1}^{\prime\prime}(% \mu)=-\frac{1}{\sigma^{2}},\qquad\mathcal{I}(\mu)=\frac{1}{\sigma^{2}}.

Since the second derivative does not depend on the data, we do not need to consider the expectation. Note that Varμ(X¯)=σ2/n=1/ℐ︀n⁢(μ)\mathop{\mathrm{Var}}\nolimits_{\mu}(\bar{X})=\sigma^{2}/n=1/\mathcal{I}_{n}(\mu) by lemma 2.2 and from lecture 13 we know that no unbiased estimator can be better.

Example 11.8.

For the exponential distribution with rate θ>0\theta>0, we have ℓ1⁢(θ)=log⁡θ−θ⁢X\ell_{1}(\theta)=\log\theta-\theta X and thus

ℓ1′⁢(θ)=1θ−X,ℓ1′′⁢(θ)=−1θ2,ℐ︀⁢(θ)=1θ2.\ell_{1}^{\prime}(\theta)=\frac{1}{\theta}-X,\qquad\ell_{1}^{\prime\prime}(% \theta)=-\frac{1}{\theta^{2}},\qquad\mathcal{I}(\theta)=\frac{1}{\theta^{2}}.

Again, the second derivative is non-random. In fact, section 11.6 shows that the dependence of the information on the scale of the parameter is a general property.

11.5 Exponential Families

All four models from section 11.4 are exponential families. For such models, the information is the second derivative of the cumulant function KK, given by definition 7.6.

Proposition 11.9.

Let f⁢(x;η)=h⁢(x)⁢exp⁡(η⁢T⁢(x)−K⁢(η))f(x;\eta)=h(x)\exp\bigl{(}\eta T(x)-K(\eta)\bigr{)} be an exponential family in its natural parametrisation. Then the Fisher information of a single observation about η\eta is given by

ℐ︀⁢(η)=K′′⁢(η)=Varη(T⁢(X)).\mathcal{I}(\eta)=K^{\prime\prime}(\eta)=\mathop{\mathrm{Var}}\nolimits_{\eta}% \bigl{(}T(X)\bigr{)}.

On the other hand, if the family is written as f⁢(x;θ)=h⁢(x)⁢exp⁡(η⁢(θ)⁢T⁢(x)−B⁢(θ))f(x;\theta)=h(x)\exp\bigl{(}\eta(\theta)T(x)-B(\theta)\bigr{)} as in definition 7.1 and if η\eta is two times continuously differentiable, then the information about θ\theta is given by

ℐ︀⁢(θ)=η′⁢(θ)2⁢K′′⁢(η⁢(θ))=η′⁢(θ)2⁢Varθ(T⁢(X)).\mathcal{I}(\theta)=\eta^{\prime}(\theta)^{2}\,K^{\prime\prime}\bigl{(}\eta(% \theta)\bigr{)}=\eta^{\prime}(\theta)^{2}\,\mathop{\mathrm{Var}}\nolimits_{% \theta}\bigl{(}T(X)\bigr{)}.
Proof.

In the natural parametrisation we have ℓ1⁢(η)=log⁡h⁢(X)+η⁢T⁢(X)−K⁢(η)\ell_{1}(\eta)=\log h(X)+\eta T(X)-K(\eta) and thus ℓ1′⁢(η)=T⁢(X)−K′⁢(η)\ell_{1}^{\prime}(\eta)=T(X)-K^{\prime}(\eta) and ℓ1′′⁢(η)=−K′′⁢(η)\ell_{1}^{\prime\prime}(\eta)=-K^{\prime\prime}(\eta). By theorem 7.7 we have 𝔼η⁢(T⁢(X))=K′⁢(η)\mathbb{E}_{\eta}\bigl{(}T(X)\bigr{)}=K^{\prime}(\eta) and Varη(T⁢(X))=K′′⁢(η)\mathop{\mathrm{Var}}\nolimits_{\eta}\bigl{(}T(X)\bigr{)}=K^{\prime\prime}(\eta) and thus, since the second derivative is non-random, we find ℐ︀⁢(η)=K′′⁢(η)=Varη(T⁢(X))\mathcal{I}(\eta)=K^{\prime\prime}(\eta)=\mathop{\mathrm{Var}}\nolimits_{\eta}% \bigl{(}T(X)\bigr{)}. For the general parametrisation we have B⁢(θ)=K⁢(η⁢(θ))B(\theta)=K\bigl{(}\eta(\theta)\bigr{)}, since both quantities are determined by the requirement that the density must integrate to one, and thus, using the chain rule, we find

ℓ1′⁢(θ)=η′⁢(θ)⁢(T⁢(X)−K′⁢(η⁢(θ))).\ell_{1}^{\prime}(\theta)=\eta^{\prime}(\theta)\,\Bigl{(}T(X)-K^{\prime}\bigl{% (}\eta(\theta)\bigr{)}\Bigr{)}.

The factor η′⁢(θ)\eta^{\prime}(\theta) is a constant and thus the variance of the score is

ℐ︀⁢(θ)=η′⁢(θ)2⁢Varθ(T⁢(X))=η′⁢(θ)2⁢K′′⁢(η⁢(θ)),\mathcal{I}(\theta)=\eta^{\prime}(\theta)^{2}\,\mathop{\mathrm{Var}}\nolimits_% {\theta}\bigl{(}T(X)\bigr{)}=\eta^{\prime}(\theta)^{2}\,K^{\prime\prime}\bigl{% (}\eta(\theta)\bigr{)},

as claimed. ∎

The proposition turns the computation of the information into that of a variance, which for the standard families can be read off tables A.1 and A.2; exercises 11.2 and 11.4 use it in this way.

11.6 Reparametrisation

The information depends on which parameter we use to describe a model, and the factor η′⁢(θ)2\eta^{\prime}(\theta)^{2} in proposition 11.9 is an instance of the general rule.

Proposition 11.10.

Let ψ=g⁢(θ)\psi=g(\theta) where g:Θ→g⁢(Θ)g\colon\Theta\to g(\Theta) is a differentiable bijection with g′⁢(θ)≠0g^{\prime}(\theta)\neq 0 for all θ∈Θ\theta\in\Theta. Let ℐ︀θ\mathcal{I}_{\theta} denote the Fisher information for a single observation when the model is parametrised by θ\theta and let ℐ︀ψ\mathcal{I}_{\psi} be the information when the model is parametrised by ψ\psi. Then,

ℐ︀ψ⁢(ψ)=ℐ︀θ⁢(θ)g′⁢(θ)2for ⁢ψ=g⁢(θ).\mathcal{I}_{\psi}(\psi)=\frac{\mathcal{I}_{\theta}(\theta)}{g^{\prime}(\theta% )^{2}}\qquad\text{for }\psi=g(\theta).
Proof.

With the new parametrisation, the density is f~⁢(x;ψ)=f⁢(x;g−1⁢(ψ))\tilde{f}(x;\psi)=f\bigl{(}x;g^{-1}(\psi)\bigr{)} and the log-likelihood for a single observation is ℓ~1⁢(ψ)=ℓ1⁢(g−1⁢(ψ))\tilde{\ell}_{1}(\psi)=\ell_{1}\bigl{(}g^{-1}(\psi)\bigr{)}. Since g′g^{\prime} never vanishes on the interval Θ\Theta, by Darboux’s theorem gg is strictly monotonic and g−1g^{-1} is differentiable. Using the chain rule and the derivative of the inverse function we get (g−1)′⁢(ψ)=1/g′⁢(θ)(g^{-1})^{\prime}(\psi)=1/g^{\prime}(\theta) for θ=g−1⁢(ψ)\theta=g^{-1}(\psi). This gives

ℓ~1′⁢(ψ)=ℓ1′⁢(θ)⁢(g−1)′⁢(ψ)=ℓ1′⁢(θ)g′⁢(θ).\tilde{\ell}_{1}^{\prime}(\psi)=\ell_{1}^{\prime}(\theta)\,(g^{-1})^{\prime}(% \psi)=\frac{\ell_{1}^{\prime}(\theta)}{g^{\prime}(\theta)}.

Since θ\theta is fixed for fixed ψ\psi, 1/g′⁢(θ)1/g^{\prime}(\theta) is a constant and taking variances under ℙθ\mathbb{P}_{\theta} gives ℐ︀ψ⁢(ψ)=ℐ︀θ⁢(θ)/g′⁢(θ)2\mathcal{I}_{\psi}(\psi)=\mathcal{I}_{\theta}(\theta)/g^{\prime}(\theta)^{2}, as claimed. ∎

For the exponential distribution, reparametrising from the rate θ\theta to the mean μ=1/θ\mu=1/\theta changes the information 1/θ21/\theta^{2} from example 11.8 to ℐ︀μ⁢(μ)=1/μ2\mathcal{I}_{\mu}(\mu)=1/\mu^{2}; see exercise 11.5. Thus, the information is not a property of the model in isolation, but depends on the parametrisation. What does not change under ψ=g⁢(θ)\psi=g(\theta) is the combination ℐ︀θ⁢(θ)⁢(d⁢θ)2\mathcal{I}_{\theta}(\theta)\,(\mathrm{d}\theta)^{2}, since d⁢ψ=g′⁢(θ)⁢d⁢θ\mathrm{d}\psi=g^{\prime}(\theta)\,\mathrm{d}\theta, and this invariance is the basis of the Jeffreys prior in lecture 26.

11.7 When the Identities Fail

The uniform distribution from example 1.2 shows what can go wrong if the conditions from section 11.1 are violated.

Example 11.11.

Let XX be uniformly distributed on [0,θ][0,\theta], i.e. f⁢(x;θ)=1/θf(x;\theta)=1/\theta for 0≤x≤θ0\leq x\leq\theta. The support of the distribution depends on θ\theta, but we can still compute the derivatives on the support to see whether the identities hold. For 0<X<θ0<X<\theta we have ℓ1⁢(θ)=−log⁡θ\ell_{1}(\theta)=-\log\theta and thus ℓ1′⁢(θ)=−1/θ\ell_{1}^{\prime}(\theta)=-1/\theta and ℓ1′′⁢(θ)=1/θ2\ell_{1}^{\prime\prime}(\theta)=1/\theta^{2}. The score is a negative constant and thus we have 𝔼θ⁢(ℓ1′⁢(θ))=−1/θ≠0\mathbb{E}_{\theta}\bigl{(}\ell_{1}^{\prime}(\theta)\bigr{)}=-1/\theta\neq 0, in contradiction to lemma 11.1. The variance is zero, whereas 𝔼θ⁢(ℓ1′⁢(θ)2)=1/θ2\mathbb{E}_{\theta}\bigl{(}\ell_{1}^{\prime}(\theta)^{2}\bigr{)}=1/\theta^{2} and −𝔼θ⁢(ℓ1′′⁢(θ))=−1/θ2-\mathbb{E}_{\theta}\bigl{(}\ell_{1}^{\prime\prime}(\theta)\bigr{)}=-1/\theta^% {2}: the three expressions for theorem 11.3 take three different values, one of which is negative. The proof breaks down at the interchange between differentiation and integration: for this model, the two sides of (11.1) with k=1k=1 differ by the term f⁢(θ;θ)=1/θf(\theta;\theta)=1/\theta from the moving upper limit, which is missed by the differentiation under the integral sign. Exercise 11.6 can be used to work out the two computations.

For this model, there is no Fisher information and the two results based on it are not applicable: the sample maximum does not satisfy the lower bound from lecture 13, the mean squared error decreases like 1/n21/n^{2} per exercise 2.2 instead of like 1/n1/n, and the asymptotic normality from lecture 17 also fails.

Summary.
  • •
    ​

    Under the conditions of section 11.1 (a support free of θ\theta, twice continuously differentiable density and differentiation under the integral sign), the score at the true parameter value has mean zero.

  • •
    ​

    The Fisher information ℐ︀n⁢(θ)\mathcal{I}_{n}(\theta) is the variance of the score, and it equals 𝔼θ⁢(ℓ′⁢(θ)2)\mathbb{E}_{\theta}\bigl{(}\ell^{\prime}(\theta)^{2}\bigr{)} and −𝔼θ⁢(ℓ′′⁢(θ))-\mathbb{E}_{\theta}\bigl{(}\ell^{\prime\prime}(\theta)\bigr{)}, the expected curvature of the log-likelihood at the truth.

  • •
    ​

    Information from independent observations can be combined, ℐ︀n⁢(θ)=n⁢ℐ︀⁢(θ)\mathcal{I}_{n}(\theta)=n\mathcal{I}(\theta); for an exponential family we have η′⁢(θ)2⁢Varθ(T)\eta^{\prime}(\theta)^{2}\mathop{\mathrm{Var}}\nolimits_{\theta}(T); under a reparametrisation ψ=g⁢(θ)\psi=g(\theta) we get ℐ︀ψ⁢(ψ)=ℐ︀θ⁢(θ)/g′⁢(θ)2\mathcal{I}_{\psi}(\psi)=\mathcal{I}_{\theta}(\theta)/g^{\prime}(\theta)^{2}.

  • •
    ​

    We have learned the following standard values: 1/(θ⁢(1−θ))1/\bigl{(}\theta(1-\theta)\bigr{)} for the Bernoulli distribution, 1/θ1/\theta for the Poisson distribution, 1/σ21/\sigma^{2} for the mean of the normal distribution and 1/θ21/\theta^{2} for the rate parameter in the exponential distribution. For the uniform distribution these relations do not hold, since the support of the distribution moves with the parameter.

  • •
    ​

    The information is the currency of the two results presented later: it gives a bound on the variance of unbiased estimators in lecture 13 and it determines the asymptotic variance of the MLE in lecture 17.

Exercise 11.1.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Bernoulli with success probability θ∈(0,1)\theta\in(0,1), and let ℓ1′⁢(θ)\ell_{1}^{\prime}(\theta) be the score of a single observation from example 11.5.

  1. 1.
    ​

    Verify lemma 11.1 for this model by computing 𝔼θ⁢(ℓ1′⁢(θ))\mathbb{E}_{\theta}\bigl{(}\ell_{1}^{\prime}(\theta)\bigr{)} directly from the two possible values of X1X_{1}.

  2. 2.
    ​

    Compute Varθ(ℓ1′⁢(θ))\mathop{\mathrm{Var}}\nolimits_{\theta}\bigl{(}\ell_{1}^{\prime}(\theta)\bigr{)} directly, and check that it agrees with the value of ℐ︀⁢(θ)\mathcal{I}(\theta) found in example 11.5.

  3. 3.
    ​

    Compare 1/ℐ︀n⁢(θ)1/\mathcal{I}_{n}(\theta) with the variance of the estimator X¯\bar{X} for θ\theta.

Exercise 11.2.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with density f⁢(x;θ)=θ−2⁢x⁢e−x/θf(x;\theta)=\theta^{-2}x\,e^{-x/\theta} for x≥0x\geq 0, where θ>0\theta>0, the gamma distribution with known shape 22 and scale θ\theta from exercise 5.2.

  1. 1.
    ​

    Compute the score ℓ1′⁢(θ)\ell_{1}^{\prime}(\theta) of a single observation and verify that it has mean zero.

  2. 2.
    ​

    Show that ℐ︀⁢(θ)=2/θ2\mathcal{I}(\theta)=2/\theta^{2}, once via −𝔼θ⁢(ℓ1′′⁢(θ))-\mathbb{E}_{\theta}\bigl{(}\ell_{1}^{\prime\prime}(\theta)\bigr{)} and once via the variance of the score, using Varθ(X1)=2⁢θ2\mathop{\mathrm{Var}}\nolimits_{\theta}(X_{1})=2\theta^{2}.

  3. 3.
    ​

    Write the density in the form of definition 7.1 with T⁢(x)=xT(x)=x, identify η⁢(θ)\eta(\theta), and recover ℐ︀⁢(θ)\mathcal{I}(\theta) from proposition 11.9.

  4. 4.
    ​

    In exercise 5.2 we found that the MLE θ^=X¯/2\hat{\theta}=\bar{X}/2 is unbiased with MSE(θ^)=θ2/(2⁢n)\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})=\theta^{2}/(2n). Compare this with 1/ℐ︀n⁢(θ)1/\mathcal{I}_{n}(\theta).

Exercise 11.3.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. from the geometric distribution with parameter p∈(0,1)p\in(0,1), i.e. f⁢(x;p)=(1−p)x−1⁢pf(x;p)=(1-p)^{x-1}p for x∈{1,2,…}x\in\{1,2,\dots\}, where 𝔼p⁢(X1)=1/p\mathbb{E}_{p}(X_{1})=1/p and Varp(X1)=(1−p)/p2\mathop{\mathrm{Var}}\nolimits_{p}(X_{1})=(1-p)/p^{2}, as in exercises 4.2 and 5.1.

  1. 1.
    ​

    Compute the score ℓ1′⁢(p)\ell_{1}^{\prime}(p) of a single observation and verify that it has mean zero.

  2. 2.
    ​

    Determine ℐ︀⁢(p)=1/(p2⁢(1−p))\mathcal{I}(p)=1/\bigl{(}p^{2}(1-p)\bigr{)} and write down ℐ︀n⁢(p)\mathcal{I}_{n}(p).

  3. 3.
    ​

    Verify the value from the previous part by taking the variance of the score.

Exercise 11.4.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Poisson with mean θ>0\theta>0.

  1. 1.
    ​

    From lecture 7 we know that the natural parameter is η=log⁡θ\eta=\log\theta and the cumulant function is K⁢(η)=eηK(\eta)=e^{\eta}. Using proposition 11.9, find the information ℐ︀η⁢(η)\mathcal{I}_{\eta}(\eta) about η\eta in terms of θ\theta.

  2. 2.
    ​

    Recover ℐ︀θ⁢(θ)=1/θ\mathcal{I}_{\theta}(\theta)=1/\theta from the previous part using proposition 11.10.

  3. 3.
    ​

    In exercise 5.7 we have estimated the probability p0=ℙθ⁢(X1=0)=e−θp_{0}=\mathbb{P}_{\theta}(X_{1}=0)=e^{-\theta}. Determine the information ℐ︀p0⁢(p0)\mathcal{I}_{p_{0}}(p_{0}) about p0p_{0} and describe its behaviour as θ→∞\theta\to\infty. Give a one-sentence answer as to why this behaviour is reasonable.

Exercise 11.5.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. exponential with rate θ>0\theta>0 and let μ=1/θ\mu=1/\theta be the mean.

  1. 1.
    ​

    Using proposition 11.10 and example 11.8, verify the value ℐ︀μ⁢(μ)=1/μ2\mathcal{I}_{\mu}(\mu)=1/\mu^{2} from section 11.6.

  2. 2.
    ​

    Show that the information of the sample about μ\mu is n/μ2n/\mu^{2} and compare the reciprocal of this quantity to the variance of the estimator X¯\bar{X} for μ\mu.

Exercise 11.6.

Let XX be uniformly distributed on the interval [0,θ][0,\theta] as described in example 11.11. Compute both sides of (11.1) for this model, using k=1k=1, and verify that the difference between the two quantities corresponds to the contribution f⁢(θ;θ)=1/θf(\theta;\theta)=1/\theta of the (moving) upper limit of the integral.