Lecture 13 The Cramer–Rao Inequality

In lecture 10 we have found the best unbiased estimators by conditioning on a complete sufficient statistic. In contrast, in this lecture we will discuss a question which could be asked about any unbiased estimator: What are the limitations of unbiased estimators? In the situation of lecture 11, we will see that the variance of an unbiased estimator cannot be smaller than 1/ℐ︀n⁢(θ)1/\mathcal{I}_{n}(\theta). This is the Cramer–Rao inequality, which gives a precise meaning to the word “optimal” and which we will use in lecture 17 to characterise the maximum likelihood estimator. In this lecture we will prove the inequality, generalise it to functions of the parameter, and then see which estimators achieve the bound. We will conclude by considering the uniform distribution, where the conditions for the bound to hold are violated and where it is possible to beat the bound.

13.1 The Inequality

Throughout this lecture, X1,…,XnX_{1},\dots,X_{n} are i.i.d. with density or probability weights f⁢(x;θ)f(x;\theta), the parameter space Θ⊆ℝ\Theta\subseteq\mathbb{R} is an open interval, and the conditions of section 11.1 hold. The score ℓ′⁢(θ)\ell^{\prime}(\theta) and the Fisher information ℐ︀n⁢(θ)=Varθ(ℓ′⁢(θ))\mathcal{I}_{n}(\theta)=\mathop{\mathrm{Var}}\nolimits_{\theta}\bigl{(}\ell^{% \prime}(\theta)\bigr{)} are as in lecture 11; we assume in addition that 0<ℐ︀n⁢(θ)<∞0<\mathcal{I}_{n}(\theta)<\infty for all θ∈Θ\theta\in\Theta, and we write f⁢(x;θ)=∏i=1nf⁢(xi;θ)f(x;\theta)=\prod_{i=1}^{n}f(x_{i};\theta) for the joint density of a value x=(x1,…,xn)x=(x_{1},\dots,x_{n}) of the sample.

The idea of the proof can be stated in advance: the covariance between an unbiased estimator WW and the score equals one, for any model, and then the Cauchy–Schwarz inequality implies that small Fisher information implies large variance. We state the result for the covariance as a separate lemma, for use later for functions of the parameter.

Lemma 13.1.

Let W=W⁢(X1,…,Xn)W=W(X_{1},\dots,X_{n}) be a statistic such that Varθ(W)<∞\mathop{\mathrm{Var}}\nolimits_{\theta}(W)<\infty for all θ∈Θ\theta\in\Theta, and assume that the expectation of WW can be differentiated under the integral sign, i.e. that

equation (13.1) (13.1)
dd⁢θ⁢∫W⁢(x)⁢f⁢(x;θ)⁢dx=∫W⁢(x)⁢∂∂θ⁢f⁢(x;θ)⁢dx,\frac{\mathrm{d}}{\mathrm{d}\theta}\int W(x)\,f(x;\theta)\,\mathrm{d}x=\int W(% x)\,\frac{\partial}{\partial\theta}f(x;\theta)\,\mathrm{d}x,

where the sum replaces the integral for discrete models, holds. Then we have

Covθ(W,ℓ′⁢(θ))=dd⁢θ⁢𝔼θ⁢(W)\mathop{\mathrm{Cov}}\nolimits_{\theta}\bigl{(}W,\ell^{\prime}(\theta)\bigr{)}% =\frac{\mathrm{d}}{\mathrm{d}\theta}\mathbb{E}_{\theta}(W)

for all θ∈Θ\theta\in\Theta.

Proof.

Since the score has mean zero (lemma 11.1), the covariance is given by the expectation of the product Covθ(W,ℓ′⁢(θ))=𝔼θ⁢(W⁢ℓ′⁢(θ))\mathop{\mathrm{Cov}}\nolimits_{\theta}\bigl{(}W,\ell^{\prime}(\theta)\bigr{)}% =\mathbb{E}_{\theta}\bigl{(}W\ell^{\prime}(\theta)\bigr{)}. As in equation (11.2), the score of the sample is the ratio ℓ′⁢(θ)=∂∂θ⁢f⁢(X;θ)/f⁢(X;θ)\ell^{\prime}(\theta)=\frac{\partial}{\partial\theta}f(X;\theta)\big{/}f(X;\theta) of the joint density and thus the density cancels when we take the expectation:

𝔼θ⁢(W⁢ℓ′⁢(θ))\displaystyle\mathbb{E}_{\theta}\bigl{(}W\ell^{\prime}(\theta)\bigr{)}
=∫W⁢(x)⁢∂∂θ⁢f⁢(x;θ)f⁢(x;θ)⁢f⁢(x;θ)⁢dx=∫W⁢(x)⁢∂∂θ⁢f⁢(x;θ)⁢dx\displaystyle=\int W(x)\,\frac{\frac{\partial}{\partial\theta}f(x;\theta)}{f(x% ;\theta)}\,f(x;\theta)\,\mathrm{d}x=\int W(x)\,\frac{\partial}{\partial\theta}% f(x;\theta)\,\mathrm{d}x
=dd⁢θ⁢∫W⁢(x)⁢f⁢(x;θ)⁢dx,\displaystyle=\frac{\mathrm{d}}{\mathrm{d}\theta}\int W(x)\,f(x;\theta)\,% \mathrm{d}x,

where the integral is over the support of the distribution, which does not depend on θ\theta, and where we used the assumption (13.1) in the last step. The final integral can be evaluated as 𝔼θ⁢(W)\mathbb{E}_{\theta}(W). This completes the proof. ∎

The assumption (13.1) is an additional exchange of differentiation and integration. For exponential families, this relation holds for all statistics with finite variance, by a similar argument from analysis as given in the proof of theorem 7.7, but we will not verify this relation for individual examples.

Theorem 13.2 (Cramer–Rao inequality).

Assume the conditions of section 11.1 and 0<ℐ︀n⁢(θ)<∞0<\mathcal{I}_{n}(\theta)<\infty. Let WW be an unbiased estimator for θ\theta with Varθ(W)<∞\mathop{\mathrm{Var}}\nolimits_{\theta}(W)<\infty which satisfies (13.1). Then we have

Varθ(W)≥1ℐ︀n⁢(θ)\mathop{\mathrm{Var}}\nolimits_{\theta}(W)\geq\frac{1}{\mathcal{I}_{n}(\theta)}

for all θ∈Θ\theta\in\Theta.

Proof.

Since WW is unbiased, we have 𝔼θ⁢(W)=θ\mathbb{E}_{\theta}(W)=\theta and using lemma 13.1 we find that the covariance Covθ(W,ℓ′⁢(θ))\mathop{\mathrm{Cov}}\nolimits_{\theta}\bigl{(}W,\ell^{\prime}(\theta)\bigr{)} is the derivative of θ\theta with respect to θ\theta, i.e. it equals 11. Using the Cauchy–Schwarz inequality, lemma A.8, for WW and ℓ′⁢(θ)\ell^{\prime}(\theta), we find

1=Covθ(W,ℓ′⁢(θ))2≤Varθ(W)⁢Varθ(ℓ′⁢(θ))=Varθ(W)⁢ℐ︀n⁢(θ),1=\mathop{\mathrm{Cov}}\nolimits_{\theta}\bigl{(}W,\ell^{\prime}(\theta)\bigr{% )}^{2}\leq\mathop{\mathrm{Var}}\nolimits_{\theta}(W)\,\mathop{\mathrm{Var}}% \nolimits_{\theta}\bigl{(}\ell^{\prime}(\theta)\bigr{)}=\mathop{\mathrm{Var}}% \nolimits_{\theta}(W)\,\mathcal{I}_{n}(\theta),

where we used definition 11.2 for the last step. Dividing by the positive number ℐ︀n⁢(θ)\mathcal{I}_{n}(\theta) gives the claim. This completes the proof. ∎

Three comments help to read the theorem. First, by proposition 11.4 the bound is 1/(n⁢ℐ︀⁢(θ))1/\bigl{(}n\mathcal{I}(\theta)\bigr{)}: no unbiased estimator can have a standard deviation which decreases faster than the 1/n1/\sqrt{n} of the sample mean (lemma 2.2). Second, the bound is independent of the estimator used, and only depends on the model: the bound states how much the data can reveal about θ\theta, regardless of the unbiased method used to estimate it. For the normal mean with known variance the bound is σ2/n=Varμ(X¯)\sigma^{2}/n=\mathop{\mathrm{Var}}\nolimits_{\mu}(\bar{X}), as shown in example 11.7, and thus X¯\bar{X} achieves this bound. Third, by theorem 2.7, the bound also acts as a lower bound for the mean squared error, but only amongst unbiased estimators: in section 2.3 we have seen a biased estimator for the normal variance, with smaller mean squared error than S2S^{2}.

13.2 Estimating a Function of the Parameter

Often the quantity of interest is a function g⁢(θ)g(\theta) of the parameter, for example the probability e−θe^{-\theta} of zero counts in example 10.2. If WW is an unbiased estimator for g⁢(θ)g(\theta), the only change in the proof is the value of the covariance.

Theorem 13.3.

Suppose the conditions of theorem 13.2 hold and that g:Θ→ℝg\colon\Theta\to\mathbb{R} is differentiable. Let WW be an unbiased estimator for g⁢(θ)g(\theta) with Varθ(W)<∞\mathop{\mathrm{Var}}\nolimits_{\theta}(W)<\infty which satisfies (13.1). Then,

Varθ(W)≥g′⁢(θ)2ℐ︀n⁢(θ)\mathop{\mathrm{Var}}\nolimits_{\theta}(W)\geq\frac{g^{\prime}(\theta)^{2}}{% \mathcal{I}_{n}(\theta)}

for all θ∈Θ\theta\in\Theta.

Proof.

We have 𝔼θ⁢(W)=g⁢(θ)\mathbb{E}_{\theta}(W)=g(\theta) and from lemma 13.1 we know that the covariance satisfies Covθ(W,ℓ′⁢(θ))=g′⁢(θ)\mathop{\mathrm{Cov}}\nolimits_{\theta}\bigl{(}W,\ell^{\prime}(\theta)\bigr{)}% =g^{\prime}(\theta). Using the Cauchy–Schwarz inequality we find g′⁢(θ)2≤Varθ(W)⁢ℐ︀n⁢(θ)g^{\prime}(\theta)^{2}\leq\mathop{\mathrm{Var}}\nolimits_{\theta}(W)\,\mathcal% {I}_{n}(\theta) and dividing by ℐ︀n⁢(θ)\mathcal{I}_{n}(\theta) completes the proof. ∎

Theorem 13.2 is the case g⁢(θ)=θg(\theta)=\theta; it is the general form which is worth memorising: identify the function gg the estimator is unbiased for, differentiate, square and divide by the information. Both numerator and denominator depend on how the model is parametrised, but by proposition 11.10 the bound itself does not: it is a property of the estimator and the model, not of the name we give to the parameter. Exercise 13.6 verifies this.

13.3 Efficiency and Attainment

We give a name to the estimators which attain the bound.

Definition 13.4.

Let WW be an unbiased estimator for g⁢(θ)g(\theta) that satisfies the assumptions of theorem 13.3. Then the efficiency of WW is given by

effθ⁡(W)=g′⁢(θ)2/ℐ︀n⁢(θ)Varθ(W),\operatorname{eff}_{\theta}(W)=\frac{g^{\prime}(\theta)^{2}/\mathcal{I}_{n}(% \theta)}{\mathop{\mathrm{Var}}\nolimits_{\theta}(W)},

where the efficiency is the ratio between the Cramer–Rao bound and the actual variance. The estimator WW is called efficient, if effθ⁡(W)=1\operatorname{eff}_{\theta}(W)=1, i.e. if Varθ(W)=g′⁢(θ)2/ℐ︀n⁢(θ)\mathop{\mathrm{Var}}\nolimits_{\theta}(W)=g^{\prime}(\theta)^{2}/\mathcal{I}_% {n}(\theta), for all θ∈Θ\theta\in\Theta.

From theorem 13.3 we know that the efficiency is contained in the interval (0,1](0,1], and since the bound is proportional to 1/n1/n, an efficiency of 1/21/2 means that the estimator uses the information from nn observations as well as an efficient estimator would use the information from n/2n/2. An efficient estimator is a UMVUE in the sense of definition 10.3, because no unbiased estimator can be better than the bound, but the converse statement is not true, since the bound is not always attainable: in exercise 13.4, the UMVUE for the rate of the exponential distribution has efficiency (n−2)/n(n-2)/n. These two results complement each other: the Lehmann–Scheffe theorem allows us to identify the best unbiased estimator, whereas the Cramer–Rao bound provides an absolute scale for comparing estimators.

The proof tells us when the bound is attained: equality in the Cramer–Rao inequality is equality in the Cauchy–Schwarz step, and this only holds for random variables which are proportional to each other.

Proposition 13.5.

With the same assumptions as in theorem 13.3, the estimator WW is efficient if and only if there is a function c⁢(θ)c(\theta) such that

equation (13.2) (13.2)
W−g⁢(θ)=c⁢(θ)⁢ℓ′⁢(θ)with probability one, for every ⁢θ∈Θ,W-g(\theta)=c(\theta)\,\ell^{\prime}(\theta)\qquad\text{with probability one, % for every }\theta\in\Theta,

and in this case c⁢(θ)=g′⁢(θ)/ℐ︀n⁢(θ)c(\theta)=g^{\prime}(\theta)/\mathcal{I}_{n}(\theta). In particular, if the model is a one-parameter exponential family as in definition 7.1 and the density is

f⁢(x;θ)=h⁢(x)⁢exp⁡(η⁢(θ)⁢T⁢(x)−B⁢(θ)),f(x;\theta)=h(x)\exp\bigl{(}\eta(\theta)T(x)-B(\theta)\bigr{)},

where η\eta is continuously differentiable with η′⁢(θ)≠0\eta^{\prime}(\theta)\neq 0 for all θ\theta, and where Tn=∑i=1nT⁢(Xi)T_{n}=\sum_{i=1}^{n}T(X_{i}) is the natural statistic, then Tn/nT_{n}/n is an efficient estimator for the mean μ⁢(θ)=𝔼θ⁢(T⁢(X1))\mu(\theta)=\mathbb{E}_{\theta}\bigl{(}T(X_{1})\bigr{)}.

Proof.

Let U=W−g⁢(θ)U=W-g(\theta) and V=ℓ′⁢(θ)V=\ell^{\prime}(\theta) be two random variables with mean zero. Then Covθ(W,V)=𝔼θ⁢(U⁢V)\mathop{\mathrm{Cov}}\nolimits_{\theta}(W,V)=\mathbb{E}_{\theta}(UV), Varθ(W)=𝔼θ⁢(U2)\mathop{\mathrm{Var}}\nolimits_{\theta}(W)=\mathbb{E}_{\theta}(U^{2}) and ℐ︀n⁢(θ)=𝔼θ⁢(V2)>0\mathcal{I}_{n}(\theta)=\mathbb{E}_{\theta}(V^{2})>0. The proof of lemma A.8 in appendix A considers the quadratic q⁢(t)=𝔼θ⁢((U−t⁢V)2)=𝔼θ⁢(U2)−2⁢t⁢𝔼θ⁢(U⁢V)+t2⁢𝔼θ⁢(V2)q(t)=\mathbb{E}_{\theta}\bigl{(}(U-tV)^{2}\bigr{)}=\mathbb{E}_{\theta}(U^{2})-% 2t\,\mathbb{E}_{\theta}(UV)+t^{2}\mathbb{E}_{\theta}(V^{2}), which is non-negative for all real tt, and the inequality states that the discriminant is non-positive. Efficiency of WW implies that the inequality is an equality, i.e. that the discriminant is zero and thus that qq has a real root t0t_{0}. Since q⁢(t0)=0q(t_{0})=0 is the expectation of the non-negative random variable (U−t0⁢V)2(U-t_{0}V)^{2}, we have U=t0⁢VU=t_{0}V with probability one, which is (13.2) with c⁢(θ)=t0c(\theta)=t_{0}. Conversely, if U=c⁢(θ)⁢VU=c(\theta)V with probability one, we have 𝔼θ⁢(U⁢V)2=c⁢(θ)2⁢𝔼θ⁢(V2)2=𝔼θ⁢(U2)⁢𝔼θ⁢(V2)\mathbb{E}_{\theta}(UV)^{2}=c(\theta)^{2}\mathbb{E}_{\theta}(V^{2})^{2}=% \mathbb{E}_{\theta}(U^{2})\,\mathbb{E}_{\theta}(V^{2}). This is an equality in the Cauchy–Schwarz inequality and thus implies efficiency. To identify c⁢(θ)c(\theta) we take the covariance of both sides of (13.2) with ℓ′⁢(θ)\ell^{\prime}(\theta). The left-hand side gives g′⁢(θ)g^{\prime}(\theta) by lemma 13.1. The right-hand side gives c⁢(θ)⁢ℐ︀n⁢(θ)c(\theta)\mathcal{I}_{n}(\theta).

For the exponential family, the proof of proposition 11.9 gives the score of a single observation as ℓi′⁢(θ)=η′⁢(θ)⁢(T⁢(Xi)−K′⁢(η⁢(θ)))\ell_{i}^{\prime}(\theta)=\eta^{\prime}(\theta)\bigl{(}T(X_{i})-K^{\prime}(% \eta(\theta))\bigr{)}, and theorem 7.7 identifies K′⁢(η⁢(θ))K^{\prime}\bigl{(}\eta(\theta)\bigr{)} with μ⁢(θ)\mu(\theta). Summing over ii we find ℓ′⁢(θ)=η′⁢(θ)⁢(Tn−n⁢μ⁢(θ))\ell^{\prime}(\theta)=\eta^{\prime}(\theta)\bigl{(}T_{n}-n\mu(\theta)\bigr{)}, and thus

Tnn−μ⁢(θ)=1n⁢η′⁢(θ)⁢ℓ′⁢(θ),\frac{T_{n}}{n}-\mu(\theta)=\frac{1}{n\,\eta^{\prime}(\theta)}\,\ell^{\prime}(% \theta),

which is (13.2) for W=Tn/nW=T_{n}/n, g=μg=\mu and c⁢(θ)=1/(n⁢η′⁢(θ))c(\theta)=1/\bigl{(}n\eta^{\prime}(\theta)\bigr{)}. The estimator Tn/nT_{n}/n is unbiased for μ⁢(θ)\mu(\theta) by construction and has finite variance K′′⁢(η⁢(θ))/nK^{\prime\prime}\bigl{(}\eta(\theta)\bigr{)}/n, and the conditions of section 11.1 hold for the family as noted there. This completes the proof. ∎

The condition (13.2) is restrictive. In an exponential family, an efficient estimator must be an affine function of the natural statistic TnT_{n}, and the coefficients of any such affine function cannot depend on θ\theta (since WW is fixed while θ\theta varies over Θ\Theta). Thus, the only functions of θ\theta which can lead to an efficient estimator are affine functions a+b⁢μ⁢(θ)a+b\,\mu(\theta) of the mean of the natural statistic, estimated by a+b⁢Tn/na+bT_{n}/n. For the Poisson distribution, an exponential family with T⁢(x)=xT(x)=x and μ⁢(θ)=θ\mu(\theta)=\theta, the sample mean is an efficient estimator for θ\theta. Exercise 13.3 shows that the attainment condition is violated in this case, by showing that e−θe^{-\theta} is not an affine function of θ\theta. Thus, the UMVUE (1−1/n)Tn(1-1/n)^{T_{n}} from example 10.6 is the best one among the unbiased estimators, but does not achieve the bound.

13.4 When the Conditions Fail

As we have seen, the proof of the inequality in this lecture involves two instances of differentiating under the integral sign, once in lemma 11.1 and once in the assumption (13.1). If the support of the distribution moves as θ\theta does, both of these exchanges will fail, since an additional term is added to the integral which is not included in the exchange.

Example 13.6.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with a uniform distribution on (0,θ)(0,\theta) and θ>0\theta>0. The support of the distribution depends on θ\theta and the conditions from section 11.1 are violated. Thus, theorem 13.2 does not apply. We can still see what the formula would be for this model: in example 11.11 the three expressions of theorem 11.3 had three different values, and taking the positive of these three values, 𝔼θ⁢(ℓ1′⁢(θ)2)=1/θ2\mathbb{E}_{\theta}\bigl{(}\ell_{1}^{\prime}(\theta)^{2}\bigr{)}=1/\theta^{2}, and multiplying by nn as in proposition 11.4, we find the formal expression n/θ2n/\theta^{2} and the formal bound θ2/n\theta^{2}/n.

Now consider the estimators for θ\theta based on the maximum M=maxi⁡XiM=\max_{i}X_{i}. The MLE MM is biased by exercise 2.2, and thus the mean squared error alone does not violate the inequality, since the inequality is a statement about unbiased estimators. The bias-corrected estimator W=(n+1)⁢M/nW=(n+1)M/n is unbiased and by exercise 10.3 has variance

Varθ(W)=θ2n⁢(n+2)<θ2nfor every ⁢n≥1,\mathop{\mathrm{Var}}\nolimits_{\theta}(W)=\frac{\theta^{2}}{n(n+2)}<\frac{% \theta^{2}}{n}\qquad\text{for every }n\geq 1,

which is of order 1/n21/n^{2} instead of 1/n1/n. An unbiased estimator with variance smaller than the formal bound in this way is said to be super-efficient. The proof fails at the covariance. The score of the sample on the support is the constant ℓ′⁢(θ)=−n/θ\ell^{\prime}(\theta)=-n/\theta, and thus we have Covθ(W,ℓ′⁢(θ))=0\mathop{\mathrm{Cov}}\nolimits_{\theta}\bigl{(}W,\ell^{\prime}(\theta)\bigr{)}=0 instead of 11: the two sides of (13.1) are dd⁢θ⁢𝔼θ⁢(W)=1\frac{\mathrm{d}}{\mathrm{d}\theta}\mathbb{E}_{\theta}(W)=1 on the left and ∫W⁢(x)⁢∂∂θ⁢f⁢(x;θ)⁢dx=−(n/θ)⁢𝔼θ⁢(W)=−n\int W(x)\frac{\partial}{\partial\theta}f(x;\theta)\,\mathrm{d}x=-(n/\theta)% \mathbb{E}_{\theta}(W)=-n on the right, and the difference is the contribution of the moving boundary of the support.

There is no contradiction: the inequality has hypotheses, and the uniform model violates these hypotheses. Super-efficiency is a property of models where the parameter is an endpoint of the support, where a single observation near the endpoint can be used to estimate θ\theta much more accurately than for a regular model. In lecture 16 we will prove the consistency of MM using a direct argument. In lecture 17 we will see that the limiting distribution of MM is not normal. In lecture 15 we will use simulation to demonstrate the 1/n21/n^{2} decay of the variance.

Summary.
  • •
    ​

    Under the conditions of section 11.1, all unbiased estimators WW for g⁢(θ)g(\theta) satisfy Varθ(W)≥g′⁢(θ)2/ℐ︀n⁢(θ)\mathop{\mathrm{Var}}\nolimits_{\theta}(W)\geq g^{\prime}(\theta)^{2}/\mathcal% {I}_{n}(\theta). For g⁢(θ)=θg(\theta)=\theta we get the bound 1/ℐ︀n⁢(θ)1/\mathcal{I}_{n}(\theta). The proof uses Covθ(W,ℓ′⁢(θ))=g′⁢(θ)\mathop{\mathrm{Cov}}\nolimits_{\theta}\bigl{(}W,\ell^{\prime}(\theta)\bigr{)}% =g^{\prime}(\theta) and the Cauchy–Schwarz inequality.

  • •
    ​

    The efficiency of an unbiased estimator is the ratio of the bound to the variance, with values in (0,1](0,1]. An efficient estimator is a UMVUE, but the converse is not true.

  • •
    ​

    The bound is saturated if and only if W−g⁢(θ)=c⁢(θ)⁢ℓ′⁢(θ)W-g(\theta)=c(\theta)\ell^{\prime}(\theta) with probability one. For the one-parameter exponential family, the natural statistic Tn/nT_{n}/n is efficient for the mean, and the only functions of θ\theta which can be used to construct an efficient estimator are affine functions of the mean.

  • •
    ​

    The bound depends on regularity conditions. For example, for the uniform distribution on (0,θ)(0,\theta), the support of the distribution depends on the parameter. The unbiased estimator (n+1)⁢M/n(n+1)M/n has variance θ2/(n⁢(n+2))\theta^{2}/\bigl{(}n(n+2)\bigr{)}, which is smaller than the formal value θ2/n\theta^{2}/n. Thus, the estimator is super-efficient.

Exercise 13.1.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. exponentially distributed with rate θ>0\theta>0. Define f⁢(x;θ)=θ⁢e−θ⁢xf(x;\theta)=\theta e^{-\theta x} for x≥0x\geq 0 and T=∑i=1nXiT=\sum_{i=1}^{n}X_{i}. Use the fact that the chi-squared distribution χk2\chi^{2}_{k} from definition A.2 coincides with the Gamma⁢(k/2,1/2)\text{Gamma}(k/2,1/2) distribution from table A.2 and that TT is Gamma⁢(n,θ)\text{Gamma}(n,\theta) distributed.

  1. 1.
    ​

    Show that Y1=2⁢θ⁢X1Y_{1}=2\theta X_{1} has density 12⁢e−y/2\tfrac{1}{2}e^{-y/2} for y≥0y\geq 0 and consequently Y1∼χ22Y_{1}\sim\chi^{2}_{2}.

  2. 2.
    ​

    Using definition A.2, show that 2⁢θ⁢T∼χ2⁢n22\theta T\sim\chi^{2}_{2n}.

  3. 3.
    ​

    Using the gamma density of TT and the change of variables formula from proposition A.19, show that you get the same result as in the previous part.

  4. 4.
    ​

    Using the mean and variance of the chi-squared distribution, find 𝔼θ⁢(T)\mathbb{E}_{\theta}(T) and Varθ(T)\mathop{\mathrm{Var}}\nolimits_{\theta}(T). Check your values by computing the moments of a single observation.

Exercise 13.2.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Bernoulli with success probability θ∈(0,1)\theta\in(0,1), and let Tn=∑i=1nXiT_{n}=\sum_{i=1}^{n}X_{i}.

  1. 1.
    ​

    Write down the Cramer–Rao bound for θ\theta, using example 11.5.

  2. 2.
    ​

    Show that X¯\bar{X} is an efficient estimator for θ\theta.

  3. 3.
    ​

    Verify condition (13.2) for W=X¯W=\bar{X} directly, by writing the score ℓ′⁢(θ)\ell^{\prime}(\theta) in terms of X¯−θ\bar{X}-\theta, and check that the resulting c⁢(θ)c(\theta) equals g′⁢(θ)/ℐ︀n⁢(θ)g^{\prime}(\theta)/\mathcal{I}_{n}(\theta).

  4. 4.
    ​

    Write down the Cramer–Rao bound for the odds g⁢(θ)=θ/(1−θ)g(\theta)=\theta/(1-\theta), and explain why no efficient estimator for the odds exists.

  5. 5.
    ​

    Show that in fact no unbiased estimator for the odds exists at all.

Exercise 13.3.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Poisson with mean θ>0\theta>0 and let Tn=∑i=1nXiT_{n}=\sum_{i=1}^{n}X_{i}.

  1. 1.
    ​

    Write down the Cramer–Rao bound for θ\theta, using example 11.6, and show that X¯\bar{X} achieves this bound.

  2. 2.
    ​

    Verify condition (13.2) for W=X¯W=\bar{X} directly, by writing the score ℓ′⁢(θ)\ell^{\prime}(\theta) in terms of X¯−θ\bar{X}-\theta, and by checking that the resulting c⁢(θ)c(\theta) equals g′⁢(θ)/ℐ︀n⁢(θ)g^{\prime}(\theta)/\mathcal{I}_{n}(\theta).

  3. 3.
    ​

    Explain why no efficient estimator for g⁢(θ)=e−θg(\theta)=e^{-\theta} exists, and why this does not contradict the fact that (1−1/n)Tn(1-1/n)^{T_{n}} is the UMVUE for e−θe^{-\theta}, using example 10.6.

Exercise 13.4.

Let X1,…,XnX_{1},\dots,X_{n}, where n≥3n\geq 3, be i.i.d. exponential with rate θ>0\theta>0, let T=∑i=1nXiT=\sum_{i=1}^{n}X_{i} and let θ~=(n−1)/T\tilde{\theta}=(n-1)/T be the unbiased version of the MLE from lecture 5.

  1. 1.
    ​

    Write down the Cramer–Rao bound for θ\theta, using example 11.8.

  2. 2.
    ​

    In exercise 5.3 we found MSE(θ~)=θ2/(n−2)\mathop{\mathrm{MSE}}\nolimits(\tilde{\theta})=\theta^{2}/(n-2). Determine the efficiency of θ~\tilde{\theta} and describe how it behaves as n→∞n\to\infty.

  3. 3.
    ​

    Show that no unbiased estimator for θ\theta can be efficient.

  4. 4.
    ​

    Show that θ~\tilde{\theta} is the UMVUE for θ\theta.

  5. 5.
    ​

    Show that X¯\bar{X} is an efficient estimator for the mean μ=1/θ\mu=1/\theta, and explain how this is compatible with the previous results.

Exercise 13.5.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. N⁢(μ,σ2)N(\mu,\sigma^{2}). Assume first that μ\mu is known and σ2>0\sigma^{2}>0 is the parameter.

  1. 1.
    ​

    Show that ℐ︀n⁢(σ2)=n/(2⁢σ4)\mathcal{I}_{n}(\sigma^{2})=n/(2\sigma^{4}) and find the Cramer–Rao bound for σ2\sigma^{2}.

  2. 2.
    ​

    Show that σ^2=1n⁢∑i=1n(Xi−μ)2\hat{\sigma}^{2}=\frac{1}{n}\sum_{i=1}^{n}(X_{i}-\mu)^{2} is an efficient estimator for σ2\sigma^{2}, using definition A.2.

  3. 3.
    ​

    Now assume that μ\mu is unknown. From lemma 2.3 we know that the sample variance S2S^{2} is an unbiased estimator for σ2\sigma^{2}. Using proposition A.3, find Var(S2)\mathop{\mathrm{Var}}\nolimits(S^{2}) and compare the result with the bound from the first part of the question. Lecture 14 shows that this bound still holds when μ\mu is unknown. What can you say about S2S^{2}, in light of example 10.7?

Exercise 13.6.

Let ψ=g⁢(θ)\psi=g(\theta) where g:Θ→g⁢(Θ)g\colon\Theta\to g(\Theta) is a continuously differentiable bijection with g′⁢(θ)≠0g^{\prime}(\theta)\neq 0 for all θ∈Θ\theta\in\Theta. Furthermore, let ℐ︀ψ\mathcal{I}_{\psi} and ℐ︀θ\mathcal{I}_{\theta} be the Fisher information for the sample with model parametrised by ψ\psi and by θ\theta, respectively.

  1. 1.
    ​

    Using propositions 11.4 and 11.10, show that ℐ︀ψ⁢(ψ)=ℐ︀θ⁢(θ)/g′⁢(θ)2\mathcal{I}_{\psi}(\psi)=\mathcal{I}_{\theta}(\theta)/g^{\prime}(\theta)^{2}.

  2. 2.
    ​

    Let WW be an unbiased estimator for g⁢(θ)g(\theta) which satisfies the assumptions of theorem 13.3. Show that the bound from theorem 13.2, for WW as an unbiased estimator for the parameter ψ\psi, coincides with the bound from theorem 13.3.