Lecture 10 Improving Estimators: Rao–Blackwell and Lehmann–Scheffe

In lecture 8 we have seen that a sufficient statistic TT can be used to summarise all the information in the data about the parameter. In this lecture we will use this idea to construct improved versions of estimators. The Rao–Blackwell theorem states that conditioning an unbiased estimator on the value of TT never increases its variance. The Lehmann–Scheffe theorem states that, if TT is a complete sufficient statistic, this procedure gives the unique best unbiased estimator. Finally, Basu’s theorem gives the independence of X¯\bar{X} and S2S^{2} for normal samples and this result will be used in lectures 25 and 29. The proofs of these three theorems all use facts about conditional expectations, as shown in propositions A.9 and A.11 in appendix A.

10.1 The Rao–Blackwell Theorem

Assume that we want to estimate a function g⁢(θ)g(\theta) of the parameter, e.g. we could be interested in the probability ℙθ⁢(X1=0)\mathbb{P}_{\theta}(X_{1}=0) of zero counts. As in lecture 2, an estimator UU is unbiased for g⁢(θ)g(\theta), if 𝔼θ⁢(U)=g⁢(θ)\mathbb{E}_{\theta}(U)=g(\theta) for all θ∈Θ\theta\in\Theta. While unbiased estimators are often easy to find, they often turn out to be wasteful, since they only use a small amount of the available sample. The Rao–Blackwell theorem uses a sufficient statistic TT to improve an existing estimator by replacing UU with the conditional expectation V=𝔼θ⁢(U|T)V=\mathbb{E}_{\theta}(U\mskip 1.0mu|\mskip 1.0muT), i.e. the average of UU for all samples which have the same value of TT as the observed sample.

Theorem 10.1 (Rao–Blackwell).

Let TT be a sufficient statistic for θ\theta and let UU be an unbiased estimator for g⁢(θ)g(\theta) with Varθ(U)<∞\mathop{\mathrm{Var}}\nolimits_{\theta}(U)<\infty for all θ\theta. Then V=𝔼θ⁢(U|T)V=\mathbb{E}_{\theta}(U\mskip 1.0mu|\mskip 1.0muT) satisfies the following properties.

  1. 1.
    ​

    The random variable VV does not depend on θ\theta, but is a function of TT alone and thus a statistic.

  2. 2.
    ​

    The statistic VV is an unbiased estimator for g⁢(θ)g(\theta).

  3. 3.
    ​

    For all θ∈Θ\theta\in\Theta we have Varθ(V)≤Varθ(U)\mathop{\mathrm{Var}}\nolimits_{\theta}(V)\leq\mathop{\mathrm{Var}}\nolimits_{% \theta}(U), with equality if and only if ℙθ⁢(U=V)=1\mathbb{P}_{\theta}(U=V)=1.

Proof.

From definition 8.1 we know that the conditional distribution of the sample (X1,…,Xn)(X_{1},\dots,X_{n}) given TT does not depend on θ\theta. Since UU is a function of the sample, the conditional distribution and thus the conditional expectation given TT do not depend on θ\theta. Thus, VV is a function of TT and by definition 1.3 a statistic.

For the second statement, using the tower property from proposition A.9, we find

𝔼θ⁢(V)=𝔼θ⁢(𝔼θ⁢(U|T))=𝔼θ⁢(U)=g⁢(θ)\mathbb{E}_{\theta}(V)=\mathbb{E}_{\theta}\bigl{(}\mathbb{E}_{\theta}(U\mskip 1% .0mu|\mskip 1.0muT)\bigr{)}=\mathbb{E}_{\theta}(U)=g(\theta)

for all θ\theta. This shows that VV is unbiased.

Finally, for the third statement, we can use the law of total variance, proposition A.11, with UU instead of XX and TT instead of YY. Since 𝔼θ⁢(U|T)=V\mathbb{E}_{\theta}(U\mskip 1.0mu|\mskip 1.0muT)=V we get

Varθ(U)=𝔼θ⁢(Varθ(U|T))+Varθ(V).\mathop{\mathrm{Var}}\nolimits_{\theta}(U)=\mathbb{E}_{\theta}\bigl{(}\mathop{% \mathrm{Var}}\nolimits_{\theta}(U\mskip 1.0mu|\mskip 1.0muT)\bigr{)}+\mathop{% \mathrm{Var}}\nolimits_{\theta}(V).

Since the conditional variance Varθ(U|T)\mathop{\mathrm{Var}}\nolimits_{\theta}(U\mskip 1.0mu|\mskip 1.0muT) is non-negative, we find Varθ(V)≤Varθ(U)\mathop{\mathrm{Var}}\nolimits_{\theta}(V)\leq\mathop{\mathrm{Var}}\nolimits_{% \theta}(U). Equality holds, if and only if 𝔼θ⁢(Varθ(U|T))=0\mathbb{E}_{\theta}\bigl{(}\mathop{\mathrm{Var}}\nolimits_{\theta}(U\mskip 1.0% mu|\mskip 1.0muT)\bigr{)}=0 and since a non-negative random variable with expectation zero is zero with probability one, this is the case if and only if Varθ(U|T)=0\mathop{\mathrm{Var}}\nolimits_{\theta}(U\mskip 1.0mu|\mskip 1.0muT)=0 with probability one, i.e. if and only if UU coincides with its conditional expectation VV with probability one. This completes the proof. ∎

Since both estimators are unbiased, the mean squared errors equal the variances, as given in theorem 2.7, and thus VV is at least as good as UU, and strictly better unless UU was already a function of TT. The procedure of passing from UU to VV is called Rao–Blackwellisation. The necessity of sufficiency is highlighted by the fact that conditioning on a statistic which is not sufficient in general will result in a quantity which depends on θ\theta but is not an estimator.

Example 10.2.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Poisson with parameter θ\theta and assume that we want to estimate g⁢(θ)=ℙθ⁢(X1=0)=e−θg(\theta)=\mathbb{P}_{\theta}(X_{1}=0)=e^{-\theta}. The indicator U=𝟏{X1=0}U=\mathbf{1}_{\{X_{1}=0\}} is unbiased for g⁢(θ)g(\theta), since 𝔼θ⁢(U)=ℙθ⁢(X1=0)\mathbb{E}_{\theta}(U)=\mathbb{P}_{\theta}(X_{1}=0). By corollary 8.3 the sum T=∑iXiT=\sum_{i}X_{i} is sufficient. Given T=tT=t, the first observation is distributed as Binomial⁢(t,1/n)\text{Binomial}(t,1/n) and thus we have

V=𝔼θ⁢(U|T)=ℙθ⁢(X1=0|T)=(1−1n)T.V=\mathbb{E}_{\theta}(U\mskip 1.0mu|\mskip 1.0muT)=\mathbb{P}_{\theta}(X_{1}=0% \mskip 1.0mu|\mskip 1.0muT)=\Bigl{(}1-\frac{1}{n}\Bigr{)}^{T}.

The Rao–Blackwellised estimator now uses all nn observations. Exercises 10.1 and 10.2 show how to find the conditional distribution and how to apply the same procedure to a continuous model.

10.2 Completeness and the Lehmann–Scheffe Theorem

The Rao–Blackwell theorem improves an existing estimator, but two different unbiased estimators could in principle be improved to two different functions of TT. The concept of completeness, definition 8.8, solves this problem: if TT is complete, there can only be one unbiased estimator for g⁢(θ)g(\theta) among the functions of TT and thus all Rao–Blackwellisations will converge to this estimator. We start our discussion by naming the target.

Definition 10.3.

An unbiased estimator VV for g⁢(θ)g(\theta) with Varθ(V)<∞\mathop{\mathrm{Var}}\nolimits_{\theta}(V)<\infty for all θ∈Θ\theta\in\Theta is a uniformly minimum variance unbiased estimator (UMVUE) for g⁢(θ)g(\theta), if Varθ(V)≤Varθ(W)\mathop{\mathrm{Var}}\nolimits_{\theta}(V)\leq\mathop{\mathrm{Var}}\nolimits_{% \theta}(W) for all θ∈Θ\theta\in\Theta and for every unbiased estimator WW for g⁢(θ)g(\theta) with finite variance.

The word “uniformly” refers to the parameter: the inequality has to hold for every θ\theta at once, and there is no reason a priori why such an estimator should exist. The following theorem shows that, whenever a complete sufficient statistic exists, a UMVUE exists as well and is easy to identify.

Theorem 10.4 (Lehmann–Scheffe).

Let TT be a complete sufficient statistic for θ\theta and let V=φ⁢(T)V=\varphi(T) be a function of TT, an unbiased estimator for g⁢(θ)g(\theta) with Varθ(V)<∞\mathop{\mathrm{Var}}\nolimits_{\theta}(V)<\infty for all θ\theta. Then VV is a UMVUE for g⁢(θ)g(\theta). Furthermore, VV is the unique UMVUE: any other unbiased estimator WW for g⁢(θ)g(\theta) with Varθ(W)=Varθ(V)\mathop{\mathrm{Var}}\nolimits_{\theta}(W)=\mathop{\mathrm{Var}}\nolimits_{% \theta}(V) for all θ\theta satisfies ℙθ⁢(W=V)=1\mathbb{P}_{\theta}(W=V)=1 for all θ\theta.

Proof.

To show the uniqueness of VV as an unbiased estimator for g⁢(θ)g(\theta) which is a function of TT, let ψ⁢(T)\psi(T) be another such function. Then we have 𝔼θ⁢(ψ⁢(T))=g⁢(θ)=𝔼θ⁢(φ⁢(T))\mathbb{E}_{\theta}\bigl{(}\psi(T)\bigr{)}=g(\theta)=\mathbb{E}_{\theta}\bigl{% (}\varphi(T)\bigr{)} for all θ\theta. The function h=ψ−φh=\psi-\varphi satisfies 𝔼θ⁢(h⁢(T))=0\mathbb{E}_{\theta}\bigl{(}h(T)\bigr{)}=0 for all θ\theta and, since TT is complete, by definition 8.8, we have ℙθ⁢(h⁢(T)=0)=1\mathbb{P}_{\theta}\bigl{(}h(T)=0\bigr{)}=1 for all θ\theta, i.e. ψ⁢(T)=φ⁢(T)\psi(T)=\varphi(T) with probability one.

Now let WW be any unbiased estimator for g⁢(θ)g(\theta) with finite variance. Then, by the Rao–Blackwell theorem, theorem 10.1, the random variable 𝔼θ⁢(W|T)\mathbb{E}_{\theta}(W\mskip 1.0mu|\mskip 1.0muT) is a function of TT which is unbiased for g⁢(θ)g(\theta). Thus, by the first step of the proof, it equals VV with probability one. By the Rao–Blackwell theorem we have then

Varθ(V)=Varθ(𝔼θ⁢(W|T))≤Varθ(W)\mathop{\mathrm{Var}}\nolimits_{\theta}(V)=\mathop{\mathrm{Var}}\nolimits_{% \theta}\bigl{(}\mathbb{E}_{\theta}(W\mskip 1.0mu|\mskip 1.0muT)\bigr{)}\leq% \mathop{\mathrm{Var}}\nolimits_{\theta}(W)

for all θ\theta, i.e. VV is a UMVUE. If Varθ(W)=Varθ(V)\mathop{\mathrm{Var}}\nolimits_{\theta}(W)=\mathop{\mathrm{Var}}\nolimits_{% \theta}(V) for all θ\theta, then the equality case of the Rao–Blackwell theorem implies W=𝔼θ⁢(W|T)=VW=\mathbb{E}_{\theta}(W\mskip 1.0mu|\mskip 1.0muT)=V with probability one. This completes the proof. ∎

The theorem turns the search for a best unbiased estimator into a two-step recipe: find a complete sufficient statistic TT, then find any function of TT which is unbiased for g⁢(θ)g(\theta), by adjusting a constant or by Rao–Blackwellising a crude unbiased estimator. For the models of this module the following two facts settle the completeness.

Proposition 10.5.
  1. 1.
    ​

    Let X1,…,XnX_{1},\dots,X_{n} be an i.i.d. sample from an exponential family of full rank, as given in definitions 7.1 and 7.10. Then the corresponding natural statistic Tn=∑i=1nT⁢(Xi)T_{n}=\sum_{i=1}^{n}T(X_{i}) is complete.

  2. 2.
    ​

    Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. and uniformly distributed on the set (0,θ)(0,\theta) where θ>0\theta>0. Then the sample maximum maxi⁡Xi\max_{i}X_{i} is complete.

This proposition is given without proof. The first statement rests on the uniqueness theorem for Laplace transforms. The second statement rests on taking derivatives of the relation ∫0θh⁢(t)⁢tn−1⁢dt=0\int_{0}^{\theta}h(t)\,t^{n-1}\,\mathrm{d}t=0 with respect to θ\theta. Both statements are allowed to be quoted. Together with corollary 8.3 and exercise 8.2 they form a complete sufficient statistic for all families of the form given in tables 1.1 and 1.2 (with mm known for the binomial). If we use a one-to-one function of a complete sufficient statistic, the new function is again complete and sufficient. Thus we can use the function X¯\bar{X} instead of ∑iXi\sum_{i}X_{i} if required.

Example 10.6.

For the Poisson sample from example 10.2, the natural parameter η=log⁡θ\eta=\log\theta takes values on all of ℝ\mathbb{R} and thus the sum T=∑iXiT=\sum_{i}X_{i} is a complete and sufficient statistic. Since X¯=T/n\bar{X}=T/n is a function of TT and is unbiased for θ\theta, it is the UMVUE for θ\theta by theorem 10.4. Similarly, the Rao–Blackwellised indicator V=(1−1/n)TV=(1-1/n)^{T} is the UMVUE for e−θe^{-\theta}.

Example 10.7.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. N⁢(μ,σ2)N(\mu,\sigma^{2}) with both parameters unknown. This is a two-parameter exponential family of full rank with natural statistic T⁢(x)=(x,x2)T(x)=(x,x^{2}), and thus the pair (X¯,S2)(\bar{X},S^{2}) is complete and sufficient; exercise 10.4 carries out the verification. Since X¯\bar{X} and S2S^{2} are unbiased for μ\mu and σ2\sigma^{2} by lemmas 2.2 and 2.3, they are the UMVUEs for these parameters.

The second example puts section 2.3 into perspective: there the biased estimator 1n+1⁢∑i(Xi−X¯)2\frac{1}{n+1}\sum_{i}(X_{i}-\bar{X})^{2} had smaller mean squared error than S2S^{2}. No contradiction, since the Lehmann–Scheffe theorem compares S2S^{2} to unbiased estimators only.

10.3 Ancillary Statistics and Basu’s Theorem

A sufficient statistic contains all the information about θ\theta. At the opposite extreme are statistics which contain none.

Definition 10.8.

A statistic A=A⁢(X1,…,Xn)A=A(X_{1},\dots,X_{n}) is ancillary for θ\theta, if the distribution of AA under ℙθ\mathbb{P}_{\theta} does not depend on θ∈Θ\theta\in\Theta.

The running example is a normal sample X1,…,Xn∼N⁢(μ,σ2)X_{1},\dots,X_{n}\sim N(\mu,\sigma^{2}) with known variance σ2\sigma^{2} and unknown mean μ\mu. Writing Zi=Xi−μZ_{i}=X_{i}-\mu we have Xi−X¯=Zi−Z¯X_{i}-\bar{X}=Z_{i}-\bar{Z} and the ZiZ_{i} are i.i.d. N⁢(0,σ2)N(0,\sigma^{2}), for every value of μ\mu. Thus, the vector of residuals (X1−X¯,…,Xn−X¯)(X_{1}-\bar{X},\dots,X_{n}-\bar{X}) is ancillary for μ\mu and every function of the residuals is ancillary, e.g. the sample variance S2S^{2} and the range maxi⁡Xi−mini⁡Xi\max_{i}X_{i}-\min_{i}X_{i}; we will consider the range in exercise 10.5.

One might expect an ancillary statistic to have nothing to do with a sufficient one. Under completeness this becomes a theorem, and it is a surprisingly effective way of proving independence.

Theorem 10.9 (Basu).

Let TT be a complete sufficient statistic for θ\theta and let AA be an ancillary statistic. Then TT and AA are independent of each other under ℙθ\mathbb{P}_{\theta} for all θ∈Θ\theta\in\Theta.

Proof.

Let BB be a set of possible values of AA. Since AA is ancillary, the probability pB=ℙθ⁢(A∈B)p_{B}=\mathbb{P}_{\theta}(A\in B) does not depend on θ\theta. We compare this probability to the conditional probability

h⁢(T)=ℙθ⁢(A∈B|T)=𝔼θ⁢(𝟏{A∈B}|T).h(T)=\mathbb{P}_{\theta}(A\in B\mskip 1.0mu|\mskip 1.0muT)=\mathbb{E}_{\theta}% \bigl{(}\mathbf{1}_{\{A\in B\}}\!\mathrel{\big{|}}\!T\bigr{)}.

Since AA is a function of the sample and since TT is sufficient, the conditional distribution of AA given TT does not depend on θ\theta. Thus, h⁢(T)h(T) is a function of TT alone. Using the tower property from proposition A.9 we find

𝔼θ⁢(h⁢(T))=ℙθ⁢(A∈B)=pB\mathbb{E}_{\theta}\bigl{(}h(T)\bigr{)}=\mathbb{P}_{\theta}(A\in B)=p_{B}

for all θ\theta. Thus the function h−pBh-p_{B} satisfies 𝔼θ⁢(h⁢(T)−pB)=0\mathbb{E}_{\theta}\bigl{(}h(T)-p_{B}\bigr{)}=0 for all θ\theta. Since TT is complete, we have h⁢(T)=pBh(T)=p_{B} with probability one for every ℙθ\mathbb{P}_{\theta}.

It remains to show that this implies independence. For any set CC of possible values of TT, the indicator function 𝟏{T∈C}\mathbf{1}_{\{T\in C\}} is a function of TT and can be taken out of the conditional expectation given TT (see proposition A.10 in appendix A). Using the tower property we find

ℙθ⁢(T∈C,A∈B)\displaystyle\mathbb{P}_{\theta}(T\in C,\ A\in B)
=𝔼θ⁢(𝟏{T∈C}⁢𝟏{A∈B})=𝔼θ⁢(𝟏{T∈C}⁢𝔼θ⁢(𝟏{A∈B}|T))\displaystyle=\mathbb{E}_{\theta}\bigl{(}\mathbf{1}_{\{T\in C\}}\mathbf{1}_{\{% A\in B\}}\bigr{)}=\mathbb{E}_{\theta}\Bigl{(}\mathbf{1}_{\{T\in C\}}\,\mathbb{% E}_{\theta}\bigl{(}\mathbf{1}_{\{A\in B\}}\!\mathrel{\big{|}}\!T\bigr{)}\Bigr{)}
=𝔼θ⁢(𝟏{T∈C}⁢pB)=ℙθ⁢(T∈C)⁢ℙθ⁢(A∈B).\displaystyle=\mathbb{E}_{\theta}\bigl{(}\mathbf{1}_{\{T\in C\}}\,p_{B}\bigr{)% }=\mathbb{P}_{\theta}(T\in C)\,\mathbb{P}_{\theta}(A\in B).

Since BB and CC were arbitrary, we have shown the independence of TT and AA and the proof is complete. ∎

The proof uses each hypothesis exactly once: ancillarity makes the unconditional probability constant in θ\theta, sufficiency makes the conditional probability a statistic, and completeness forces these two to be the same. The most important application of this result in this module is the following.

Corollary 10.10.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. N⁢(μ,σ2)N(\mu,\sigma^{2}). Then X¯\bar{X} and S2S^{2} are independent.

Proof.

Fixing the value of σ2\sigma^{2}, we can consider μ∈ℝ\mu\in\mathbb{R} to be the only parameter. The density of a single observation is

f⁢(x;μ)=12⁢π⁢σ2⁢e−x2/(2⁢σ2)⁢exp⁡(μσ2⁢x−μ22⁢σ2).f(x;\mu)=\frac{1}{\sqrt{2\pi\sigma^{2}}}\,e^{-x^{2}/(2\sigma^{2})}\exp\Bigl{(}% \frac{\mu}{\sigma^{2}}\,x-\frac{\mu^{2}}{2\sigma^{2}}\Bigr{)}.

This is of the form given in definition 7.1 with natural statistic T⁢(x)=xT(x)=x and natural parameter η⁢(μ)=μ/σ2\eta(\mu)=\mu/\sigma^{2}. As μ\mu ranges over ℝ\mathbb{R}, so does η⁢(μ)=μ/σ2\eta(\mu)=\mu/\sigma^{2}, that is, the natural parameter takes all values in ℝ\mathbb{R}, and thus the family is of full rank. By corollary 8.3 and proposition 10.5, the sum ∑iXi\sum_{i}X_{i}, and thus X¯\bar{X}, is a complete and sufficient statistic for μ\mu. Furthermore, S2S^{2} is a function of the residuals Xi−X¯X_{i}-\bar{X} and thus, as shown above, is ancillary for μ\mu. By Basu’s theorem, theorem 10.9, the statistics X¯\bar{X} and S2S^{2} are independent under ℙμ,σ2\mathbb{P}_{\mu,\sigma^{2}} for every μ\mu and, since σ2\sigma^{2} was arbitrary, the claim is proved. ∎

This completes the proof of the independence statement of proposition A.3. The result from lectures 25 and 29 on the ratio of X¯−μ\bar{X}-\mu to SS uses this proposition. The proof contains a useful trick: when both parameters are unknown, neither X¯\bar{X} nor S2S^{2} can be ancillary and we cannot apply Basu’s theorem directly. The solution to this problem is to fix the value of σ2\sigma^{2}, i.e. to consider the independence statement for one distribution at a time.

In lecture 13 we will see a second way to derive optimality, using the Cramer–Rao bound, which does not require conditioning.

Summary.
  • •
    ​

    Rao–Blackwell: If an unbiased estimator is conditioned on a sufficient statistic TT, the conditioned estimator is unbiased and has no larger variance.

  • •
    ​

    A UMVUE is an unbiased estimator which has minimal variance among all unbiased estimators, for all θ\theta.

  • •
    ​

    Lehmann–Scheffe: If TT is complete and sufficient, then the only function of TT which is unbiased for g⁢(θ)g(\theta) is the UMVUE for g⁢(θ)g(\theta).

  • •
    ​

    The natural statistic for a full-rank exponential family and the maximum of a uniform sample are complete, but without proof.

  • •
    ​

    Basu: A complete sufficient statistic is independent of all ancillary statistics. Using this result, one can show that X¯\bar{X} and S2S^{2} of a normal sample are independent.

Exercise 10.1.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Poisson with parameter θ>0\theta>0, let T=∑i=1nXiT=\sum_{i=1}^{n}X_{i} and let U=𝟏{X1=0}U=\mathbf{1}_{\{X_{1}=0\}}, as in example 10.2. You may use the fact that a sum of independent Poisson random variables is again Poisson, with the parameters added.

  1. 1.
    ​

    Show that UU is an unbiased estimator for e−θe^{-\theta} and determine the variance of this estimator.

  2. 2.
    ​

    Show that for t∈{0,1,2,…}t\in\{0,1,2,\dots\} and x∈{0,1,…,t}x\in\{0,1,\dots,t\} we have

    ℙθ⁢(X1=x|T=t)=(tx)⁢(1n)x⁢(1−1n)t−x,\mathbb{P}_{\theta}(X_{1}=x\mskip 1.0mu|\mskip 1.0muT=t)=\binom{t}{x}\Bigl{(}% \frac{1}{n}\Bigr{)}^{x}\Bigl{(}1-\frac{1}{n}\Bigr{)}^{t-x},

    i.e. that the conditional distribution of X1X_{1} given T=tT=t is Binomial⁢(t,1/n)\text{Binomial}(t,1/n) and that this distribution does not depend on θ\theta.

  3. 3.
    ​

    Deduce that V=𝔼θ⁢(U|T)=(1−1/n)TV=\mathbb{E}_{\theta}(U\mskip 1.0mu|\mskip 1.0muT)=(1-1/n)^{T}.

  4. 4.
    ​

    Show directly that VV is an unbiased estimator for e−θe^{-\theta}, using the probability generating function 𝔼θ⁢(sT)=exp⁡(n⁢θ⁢(s−1))\mathbb{E}_{\theta}(s^{T})=\exp\bigl{(}n\theta(s-1)\bigr{)} of the Poisson distribution with parameter n⁢θn\theta.

  5. 5.
    ​

    Explain why VV is the UMVUE for e−θe^{-\theta} and discuss the differences between this UMVUE and the MLE e−X¯e^{-\bar{X}}, introduced in exercise 5.7.

Exercise 10.2.

Let n≥2n\geq 2 and let X1,…,XnX_{1},\dots,X_{n} be i.i.d. exponential with rate θ>0\theta>0, i.e. f⁢(x;θ)=θ⁢e−θ⁢xf(x;\theta)=\theta e^{-\theta x} for x≥0x\geq 0 and let T=∑i=1nXiT=\sum_{i=1}^{n}X_{i}. You may use the fact that the sum of kk independent random variables of this type has density fk⁢(t)=θk⁢tk−1⁢e−θ⁢t/(k−1)!f_{k}(t)=\theta^{k}t^{k-1}e^{-\theta t}/(k-1)! for t≥0t\geq 0.

  1. 1.
    ​

    Explain why X1X_{1} is an unbiased estimator for the mean 1/θ1/\theta and why TT is sufficient for θ\theta.

  2. 2.
    ​

    Let R=∑i=2nXiR=\sum_{i=2}^{n}X_{i}. Use the independence of X1X_{1} and RR to find the joint density of (X1,T)(X_{1},T). Deduce the conditional density of X1X_{1} given T=sT=s. The result is

    f⁢(x|s)=(n−1)⁢(s−x)n−2sn−1for ⁢0<x<s,f(x\mskip 1.0mu|\mskip 1.0mus)=\frac{(n-1)(s-x)^{n-2}}{s^{n-1}}\qquad\text{for% }0<x<s,

    and zero elsewhere. Check that this density does not depend on θ\theta.

  3. 3.
    ​

    Show that 𝔼θ⁢(X1|T=s)=s/n\mathbb{E}_{\theta}(X_{1}\mskip 1.0mu|\mskip 1.0muT=s)=s/n and conclude that Rao–Blackwellising X1X_{1} leads to the sample mean X¯\bar{X}, again unbiased for 1/θ1/\theta.

  4. 4.
    ​

    Find a second argument for 𝔼θ⁢(X1|T)=T/n\mathbb{E}_{\theta}(X_{1}\mskip 1.0mu|\mskip 1.0muT)=T/n, based on the symmetry of the sample, without any integration.

Exercise 10.3.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. uniformly distributed on the set (0,θ)(0,\theta) where θ>0\theta>0, and let M=maxi⁡XiM=\max_{i}X_{i}.

  1. 1.
    ​

    Using the mean of MM from exercise 2.2, show that θ^=(n+1)⁢M/n\hat{\theta}=(n+1)M/n is an unbiased estimator for θ\theta.

  2. 2.
    ​

    Show that θ^\hat{\theta} is the UMVUE for θ\theta.

  3. 3.
    ​

    The estimator θ~=2⁢X¯\tilde{\theta}=2\bar{X} is also unbiased for θ\theta. Determine the variances of θ^\hat{\theta} and of θ~\tilde{\theta} and verify that they are ordered as in theorem 10.4.

  4. 4.
    ​

    Without computing any conditional expectations, find 𝔼θ⁢(θ~|M)\mathbb{E}_{\theta}(\tilde{\theta}\mskip 1.0mu|\mskip 1.0muM).

Exercise 10.4.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. N⁢(μ,σ2)N(\mu,\sigma^{2}), where both parameters are unknown. This is the situation of example 10.7.

  1. 1.
    ​

    Write the density of an observation in the form from definition 7.8 and identify the natural statistic T⁢(x)=(x,x2)T(x)=(x,x^{2}) and the natural parameter η⁢(μ,σ2)\eta(\mu,\sigma^{2}).

  2. 2.
    ​

    Show that the family of densities has full rank, in the sense of definition 7.10, by determining the range of η\eta.

  3. 3.
    ​

    Deduce that (∑iXi,∑iXi2)\bigl{(}\sum_{i}X_{i},\sum_{i}X_{i}^{2}\bigr{)} is complete and sufficient and argue that the same holds for the pair (X¯,S2)(\bar{X},S^{2}).

Exercise 10.5.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. N⁢(μ,σ2)N(\mu,\sigma^{2}) where σ2\sigma^{2} is known and μ∈ℝ\mu\in\mathbb{R} is unknown. Furthermore, let R=maxi⁡Xi−mini⁡XiR=\max_{i}X_{i}-\min_{i}X_{i} be the sample range.

  1. 1.
    ​

    Show that RR is ancillary for μ\mu.

  2. 2.
    ​

    Show that X¯\bar{X} and RR are independent.

  3. 3.
    ​

    Now assume that μ\mu is known and σ2>0\sigma^{2}>0 is unknown. Is RR ancillary for σ2\sigma^{2}? Is R/σR/\sigma a statistic?