Lecture 8 Sufficiency

In example 4.2 we have observed ten coin tosses and got seven heads. The likelihood in this example is L⁢(θ)=θ7⁢(1−θ)3L(\theta)=\theta^{7}(1-\theta)^{3}. This function depends on the data only via the number 77 of heads, but not on the order of heads and tails, i.e. for learning about θ\theta we can replace the full record of observations by this single number. The idea of this example is made precise in this lecture: A sufficient statistic contains all information about the parameter. The factorisation theorem helps to find sufficient statistics. The concept of sufficiency forms the basis of improvements of estimators in lecture 10.

In this lecture we write X=(X1,…,Xn)X=(X_{1},\dots,X_{n}) for the data, x=(x1,…,xn)x=(x_{1},\dots,x_{n}) for an observed value, and

f⁢(x;θ)=∏i=1nf⁢(xi;θ)f(x;\theta)=\prod_{i=1}^{n}f(x_{i};\theta)

for the joint density or joint probability weights of the sample. The vector argument indicates whether we consider the joint function of the sample, or the function ff of a single observation. This expression, as a function of θ\theta for fixed data, is the likelihood L⁢(θ)L(\theta) from lecture 4.

8.1 Sufficient Statistics

A statistic T=T⁢(X1,…,Xn)T=T(X_{1},\dots,X_{n}) in the sense of definition 1.3 reduces the data to a smaller summary, and usually some information is lost in the process. The following definition expresses that no information about θ\theta is lost.

Definition 8.1.

A statistic T=T⁢(X1,…,Xn)T=T(X_{1},\dots,X_{n}) is sufficient for θ\theta, if the conditional distribution of (X1,…,Xn)(X_{1},\dots,X_{n}) given the value T=tT=t does not depend on θ\theta, for all values tt of TT.

If TT is sufficient, anyone who knows the value tt of TT can generate a new data set with the same distribution as the original data set, by sampling from the conditional distribution of XX given T=tT=t. Since this distribution does not depend on θ\theta, all information about θ\theta that we could obtain by studying the data set can now be obtained by studying tt alone.

For the coin tosses in the introduction, exercise 8.1 shows that the definition is correct: if the number of heads T=∑i=1nXiT=\sum_{i=1}^{n}X_{i} is given as tt, each of the (nt)\binom{n}{t} possible arrangements of heads and tails has conditional probability 1/(nt)1/\binom{n}{t}, independent of the value of θ\theta. Thus, the positions of the heads do not contain any information about the coin and TT is sufficient.

Two observations help to place the definition. First, the whole sample T=(X1,…,Xn)T=(X_{1},\dots,X_{n}) is always sufficient, since given the whole sample there is no variability left; sufficiency alone therefore says nothing about how much the data have been reduced, and we return to this in section 8.4. Secondly, a one-to-one function φ⁢(T)\varphi(T) of a sufficient statistic carries the same information and is sufficient as well; for the coin tosses the sample mean X¯=T/n\bar{X}=T/n is thus sufficient too.

8.2 The Factorisation Theorem

Directly verifying definition 8.1 requires knowledge of the distribution of TT and involves a conditional probability which, for most models, is tedious to compute. The following theorem replaces the computation with an inspection of the joint density.

Theorem 8.2 (Fisher–Neyman factorisation theorem).

The statistic TT is sufficient for θ\theta, if and only if there are functions gg and hh such that

equation (8.1) (8.1)
f⁢(x;θ)=g⁢(T⁢(x);θ)⁢h⁢(x)f(x;\theta)=g\bigl{(}T(x);\theta\bigr{)}\,h(x)

for all data vectors xx and all θ∈Θ\theta\in\Theta.

The function gg may depend on the data only through T⁢(x)T(x), and the function hh may not depend on θ\theta at all: the parameter and the data interact only through TT. Nothing in the statement requires TT or θ\theta to be one-dimensional. We give a complete proof for discrete data and then explain what changes in the continuous case.

Proof for discrete data.

Since we have discrete observations, the joint probability weight f⁢(x;θ)=ℙθ⁢(X=x)f(x;\theta)=\mathbb{P}_{\theta}(X=x) is the probability of observing the data vector xx.

Assume first that TT is sufficient and let xx be a data vector with t=T⁢(x)t=T(x). Since the event {X=x}\{X=x\} is contained in the event {T=t}\{T=t\}, we have

ℙθ⁢(X=x)=ℙθ⁢(X=x,T=t)=ℙθ⁢(T=t)⁢ℙθ⁢(X=x|T=t)\mathbb{P}_{\theta}(X=x)=\mathbb{P}_{\theta}(X=x,\,T=t)=\mathbb{P}_{\theta}(T=% t)\,\mathbb{P}_{\theta}(X=x\mskip 1.0mu|\mskip 1.0muT=t)

whenever ℙθ⁢(T=t)>0\mathbb{P}_{\theta}(T=t)>0. Using the sufficiency of the statistic, the conditional probability on the right does not depend on θ\theta and thus is a function h⁢(x)h(x) of the data. The first factor is a function g⁢(t;θ)g(t;\theta) of tt and θ\theta. If ℙθ⁢(T=t)=0\mathbb{P}_{\theta}(T=t)=0 for the given θ\theta, both sides in (8.1) equal zero since g⁢(t;θ)=0g(t;\theta)=0. For the data vectors xx with ℙθ⁢(T=T⁢(x))=0\mathbb{P}_{\theta}(T=T(x))=0 for every θ\theta, which are never observed, we define h⁢(x)=0h(x)=0. Thus (8.1) holds for all xx and θ\theta.

Conversely, assume that (8.1) holds and let tt be a value of TT and θ\theta a parameter value with ℙθ⁢(T=t)>0\mathbb{P}_{\theta}(T=t)>0. Summing the factorisations for all data vectors yy with T⁢(y)=tT(y)=t, we find

ℙθ⁢(T=t)=∑y:T⁢(y)=tf⁢(y;θ)=g⁢(t;θ)⁢∑y:T⁢(y)=th⁢(y)=g⁢(t;θ)⁢H⁢(t),\mathbb{P}_{\theta}(T=t)=\sum_{y\colon T(y)=t}f(y;\theta)=g(t;\theta)\sum_{y% \colon T(y)=t}h(y)=g(t;\theta)\,H(t),

where H⁢(t)H(t) is the last sum which does not depend on θ\theta. Since the left-hand side is positive, we have g⁢(t;θ)≠0g(t;\theta)\neq 0. For data xx with T⁢(x)=tT(x)=t we get

ℙθ⁢(X=x|T=t)=ℙθ⁢(X=x)ℙθ⁢(T=t)=g⁢(t;θ)⁢h⁢(x)g⁢(t;θ)⁢H⁢(t)=h⁢(x)H⁢(t),\mathbb{P}_{\theta}(X=x\mskip 1.0mu|\mskip 1.0muT=t)=\frac{\mathbb{P}_{\theta}% (X=x)}{\mathbb{P}_{\theta}(T=t)}=\frac{g(t;\theta)\,h(x)}{g(t;\theta)\,H(t)}=% \frac{h(x)}{H(t)},

and for xx with T⁢(x)≠tT(x)\neq t the conditional probability is zero. Thus, in both cases the value does not depend on θ\theta and TT is sufficient. This completes the proof. ∎

For continuous data the event {T=t}\{T=t\} has probability zero and the conditional distribution of XX given T=tT=t cannot be described as a quotient of probabilities anymore. If the statistic can be extended to a change of variables, i.e. if we can find coordinates UU such that x↦(T⁢(x),U⁢(x))x\mapsto\bigl{(}T(x),U(x)\bigr{)} is a bijection, then the argument presented above can be applied with densities instead of probabilities: The joint density of (T,U)(T,U) will factor in a similar way, and after integrating out uu and dividing we will obtain a conditional density of UU given T=tT=t which does not depend on θ\theta. In general, no such completion of the proof is possible, and the proof will have to replace the family of measures by one dominating probability measure, for example a weighted sum of countably many of the ℙθ\mathbb{P}_{\theta}, such that the density of ℙθ\mathbb{P}_{\theta} is a function of TT. We will not go into detail here and will use theorem 8.2 for continuous models in analogy to the proof for discrete models, with the understanding that the factorisation, similar to the density, only needs to hold up to sets of probability zero.

The factorisation has a consequence for maximum likelihood. For fixed data the likelihood is L⁢(θ)=g⁢(T⁢(x);θ)⁢h⁢(x)L(\theta)=g\bigl{(}T(x);\theta\bigr{)}h(x), and the factor h⁢(x)h(x) does not affect where the maximum over θ\theta is attained: the maximum likelihood estimator, whenever it is unique, is a function of any sufficient statistic. The examples of lecture 5 confirm this: the Bernoulli MLE is T/nT/n, and the uniform MLE is the sample maximum, which exercise 8.2 shows to be sufficient.

8.3 Sufficiency in Exponential Families

For the exponential families from lecture 7, no work is needed to find a sufficient statistic, since the densities are already given in factorised form.

Corollary 8.3.

Let X1,…,XnX_{1},\dots,X_{n} be an i.i.d. sample from a one-parameter exponential family as given in definition 7.1 and let TT be the natural statistic. Then Tn=∑i=1nT⁢(Xi)T_{n}=\sum_{i=1}^{n}T(X_{i}) is sufficient for θ\theta.

Proof.

From lemma 7.5 we know that the joint density of the sample is

f⁢(x;θ)=(∏i=1nh⁢(xi))⁢exp⁡(η⁢(θ)⁢∑i=1nT⁢(xi)−n⁢B⁢(θ)).f(x;\theta)=\Bigl{(}\prod_{i=1}^{n}h(x_{i})\Bigr{)}\exp\Bigl{(}\eta(\theta)% \sum_{i=1}^{n}T(x_{i})-nB(\theta)\Bigr{)}.

Thus, we can write the joint density in the form (8.1) with g⁢(t;θ)=exp⁡(η⁢(θ)⁢t−n⁢B⁢(θ))g(t;\theta)=\exp\bigl{(}\eta(\theta)t-nB(\theta)\bigr{)} and h⁢(x)=∏ih⁢(xi)h(x)=\prod_{i}h(x_{i}). Consequently, TnT_{n} is sufficient by theorem 8.2. ∎

The same argument shows that for the kk-parameter exponential families of lecture 7 the vector (∑iT1⁢(Xi),…,∑iTk⁢(Xi))\bigl{(}\sum_{i}T_{1}(X_{i}),\dots,\sum_{i}T_{k}(X_{i})\bigr{)} of summed natural statistics is sufficient for the parameter vector. It is instructive to see the factorisation directly in the standard models.

Example 8.4.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Poisson with parameter θ>0\theta>0. Then the joint probability weights are given by

f⁢(x;θ)=∏i=1nθxixi!⁢e−θ=θ∑ixi⁢e−n⁢θ⋅1∏ixi!,f(x;\theta)=\prod_{i=1}^{n}\frac{\theta^{x_{i}}}{x_{i}!}\,e^{-\theta}=\theta^{% \sum_{i}x_{i}}\,e^{-n\theta}\cdot\frac{1}{\prod_{i}x_{i}!},

which is of the form (8.1) with g⁢(t;θ)=θt⁢e−n⁢θg(t;\theta)=\theta^{t}e^{-n\theta} and h⁢(x)=1/∏ixi!h(x)=1/\prod_{i}x_{i}!. Thus the sum T=∑iXiT=\sum_{i}X_{i} is sufficient for θ\theta. Similarly, the sample mean X¯\bar{X} is sufficient. For data sets with the same sum, the likelihood functions are proportional: this was already shown in exercise 4.1 for the log-likelihood, and lecture 9 illustrates this in R.

For the normal distribution N⁢(μ,σ2)N(\mu,\sigma^{2}), where both parameters are unknown, we can find the sufficient statistic by expanding the squares in the exponent of the density. It transpires that the joint density depends on the data only via the pair T=(∑iXi,∑iXi2)T=\bigl{(}\sum_{i}X_{i},\sum_{i}X_{i}^{2}\bigr{)}. Thus, TT is a sufficient statistic for (μ,σ2)(\mu,\sigma^{2}). Since ∑i(xi−x¯)2=∑ixi2−n⁢x¯2\sum_{i}(x_{i}-\bar{x})^{2}=\sum_{i}x_{i}^{2}-n\bar{x}^{2}, the pair (X¯,S2)(\bar{X},S^{2}) maps to TT one-to-one and thus is also a sufficient statistic. This is the pair of statistics on which the maximum likelihood estimator from example 5.4 is based. If σ2\sigma^{2} is known, the factor exp⁡(−∑ixi2/(2⁢σ2))\exp\bigl{(}-\sum_{i}x_{i}^{2}/(2\sigma^{2})\bigr{)} can be moved into h⁢(x)h(x) and then X¯\bar{X} alone suffices as a sufficient statistic for μ\mu. The factorisation is worked out in exercise 8.4.

8.4 Minimal Sufficiency and Completeness

As we have seen in section 8.1, the whole sample is always sufficient. For an i.i.d. sample, the vector of ordered observations is also sufficient, since the joint density does not change when the xix_{i} are interchanged. The aim of this section is to characterise the statistic which allows us to reduce the data as much as possible, while still retaining all the information.

Definition 8.5.

A sufficient statistic TT is minimal sufficient for θ\theta, if for every sufficient statistic SS there is a function φ\varphi such that T=φ⁢(S)T=\varphi(S).

A minimal sufficient statistic can thus be computed from every other sufficient statistic. The definition is awkward to check directly, because it refers to all sufficient statistics at once; the following criterion instead compares the likelihood functions of two data sets.

Theorem 8.6.

Let TT be a statistic, such that every data vector has positive likelihood for some θ∈Θ\theta\in\Theta, and such that for all data vectors xx and yy we have

f⁢(x;θ)f⁢(y;θ)⁢ does not depend on ⁢θif and only ifT⁢(x)=T⁢(y).\frac{f(x;\theta)}{f(y;\theta)}\ \text{ does not depend on }\theta\qquad\text{% if and only if}\qquad T(x)=T(y).

Then TT is minimal sufficient for θ\theta.

The proof of this theorem is only sketched here. The statement of sufficiency follows from theorem 8.2: For every value tt of TT we can choose a fixed data vector xtx_{t} such that T⁢(xt)=tT(x_{t})=t. Then the ratio f⁢(x;θ)/f⁢(xT⁢(x);θ)f(x;\theta)/f(x_{T(x)};\theta), by assumption, does not depend on θ\theta and can be used as h⁢(x)h(x) while g⁢(t;θ)=f⁢(xt;θ)g(t;\theta)=f(x_{t};\theta). To see minimality, let SS be any sufficient statistic. Then we can write ff as f⁢(x;θ)=g~⁢(S⁢(x);θ)⁢h~⁢(x)f(x;\theta)=\tilde{g}(S(x);\theta)\,\tilde{h}(x). If S⁢(x)=S⁢(y)S(x)=S(y), then f⁢(x;θ)/f⁢(y;θ)=h~⁢(x)/h~⁢(y)f(x;\theta)/f(y;\theta)=\tilde{h}(x)/\tilde{h}(y) does not depend on θ\theta and thus we have T⁢(x)=T⁢(y)T(x)=T(y) by assumption: the value of TT is already determined by the value of SS and TT is a function of SS.

The criterion states that T⁢(x)T(x) contains the likelihood function of the data up to a constant factor. We apply the criterion once, to the Poisson sum. The two-parameter normal case is worked out in exercise 8.4.

Example 8.7.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Poisson with parameter θ>0\theta>0, and let T=∑iXiT=\sum_{i}X_{i}. For two data vectors xx and yy, the joint probability weights of example 8.4 give

f⁢(x;θ)f⁢(y;θ)=θ∑ixi⁢e−n⁢θ/∏ixi!θ∑iyi⁢e−n⁢θ/∏iyi!=θ∑ixi−∑iyi⁢∏iyi!∏ixi!.\frac{f(x;\theta)}{f(y;\theta)}=\frac{\theta^{\sum_{i}x_{i}}e^{-n\theta}/\prod% _{i}x_{i}!}{\theta^{\sum_{i}y_{i}}e^{-n\theta}/\prod_{i}y_{i}!}=\theta^{\sum_{% i}x_{i}-\sum_{i}y_{i}}\,\frac{\prod_{i}y_{i}!}{\prod_{i}x_{i}!}.

If ∑ixi=∑iyi\sum_{i}x_{i}=\sum_{i}y_{i}, the power of θ\theta disappears and the ratio is a constant. If the sums differ by k≠0k\neq 0, the ratio contains the factor θk\theta^{k}, which is not constant as θ\theta ranges over (0,∞)(0,\infty). Thus the ratio is free of θ\theta if and only if ∑ixi=∑iyi\sum_{i}x_{i}=\sum_{i}y_{i}, and TT is minimal sufficient by theorem 8.6.

For the results of lecture 10 we will need one more property of a statistic, concerning functions of the statistic which are unbiased estimators for zero.

Definition 8.8.

A statistic TT is complete, if for every function gg with finite mean g⁢(T)g(T) under ℙθ\mathbb{P}_{\theta} for all θ∈Θ\theta\in\Theta, the condition 𝔼θ⁢(g⁢(T))=0\mathbb{E}_{\theta}\bigl{(}g(T)\bigr{)}=0 for all θ∈Θ\theta\in\Theta implies ℙθ⁢(g⁢(T)=0)=1\mathbb{P}_{\theta}\bigl{(}g(T)=0\bigr{)}=1 for all θ∈Θ\theta\in\Theta.

This means that the only function of a complete statistic which has expectation zero for all parameter values is the zero function. The condition is violated for statistics which do not reduce the data enough: in exercise 8.6 we will see that the whole Bernoulli sample is sufficient, but not complete, whereas the total ∑iXi\sum_{i}X_{i} is complete, as we will see in lecture 10 (proposition 10.5). The property is important because any two unbiased estimators for the same quantity, which are both functions of a complete statistic TT, differ by a function of TT with expectation zero, and by completeness this function must be zero with probability one: for each quantity there is at most one unbiased estimator which is a function of a complete statistic. In lecture 10 we will use this uniqueness to turn the Rao–Blackwell improvement into an optimality result. In the lecture we will also discuss which of the standard statistics are complete.

8.5 The Sample Maximum and Minimum

So far we have considered sufficient statistics which are sums, and we have seen the distributions of these sums in appendix A. For the uniform distribution we will consider the sample maximum, and in order to determine the bias or the mean squared error of the maximum, we need to know the distribution of the maximum. The following result gives the distribution of the maximum and minimum of an i.i.d. sample.

Proposition 8.9.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with distribution function FF and density ff. Furthermore, let X(n)=maxi⁡XiX_{(n)}=\max_{i}X_{i} and X(1)=mini⁡XiX_{(1)}=\min_{i}X_{i}. Then the random variables X(n)X_{(n)} and X(1)X_{(1)} have densities

f(n)⁢(y)=n⁢F⁢(y)n−1⁢f⁢(y)andf(1)⁢(y)=n⁢(1−F⁢(y))n−1⁢f⁢(y).f_{(n)}(y)=n\,F(y)^{n-1}f(y)\qquad\text{and}\qquad f_{(1)}(y)=n\,\bigl{(}1-F(y% )\bigr{)}^{n-1}f(y).
Proof.

The maximum is at most yy, if and only if all observations are at most yy. Using independence we get

ℙ⁢(X(n)≤y)=ℙ⁢(X1≤y,…,Xn≤y)=∏i=1nℙ⁢(Xi≤y)=F⁢(y)n.\mathbb{P}(X_{(n)}\leq y)=\mathbb{P}(X_{1}\leq y,\dots,X_{n}\leq y)=\prod_{i=1% }^{n}\mathbb{P}(X_{i}\leq y)=F(y)^{n}.

Using the chain rule for differentiation with respect to yy, we find that the maximum has density n⁢F⁢(y)n−1⁢f⁢(y)nF(y)^{n-1}f(y). Similarly, the minimum is strictly greater than yy, if and only if all observations are strictly greater than yy. Thus we get

ℙ⁢(X(1)≤y)=1−ℙ⁢(X1>y,…,Xn>y)=1−(1−F⁢(y))n.\mathbb{P}(X_{(1)}\leq y)=1-\mathbb{P}(X_{1}>y,\dots,X_{n}>y)=1-\bigl{(}1-F(y)% \bigr{)}^{n}.

Using the chain rule for differentiation we find n⁢(1−F⁢(y))n−1⁢f⁢(y)n\bigl{(}1-F(y)\bigr{)}^{n-1}f(y). This completes the proof. ∎

Example 8.10.

Let X1,…,Xn∼Uniform⁢(0,θ)X_{1},\dots,X_{n}\sim\text{Uniform}(0,\theta) be i.i.d., as in example 1.2. The distribution function is F⁢(y)=y/θF(y)=y/\theta for 0≤y≤θ0\leq y\leq\theta, and the density is f⁢(y)=1/θf(y)=1/\theta on [0,θ][0,\theta]. By proposition 8.9 the sample maximum has density

f(n)⁢(y)=n⁢(yθ)n−1⁢1θ=n⁢yn−1θn,0≤y≤θ,f_{(n)}(y)=n\Bigl{(}\frac{y}{\theta}\Bigr{)}^{n-1}\frac{1}{\theta}=\frac{n\,y^% {n-1}}{\theta^{n}},\qquad 0\leq y\leq\theta,

and f(n)⁢(y)=0f_{(n)}(y)=0 outside this interval. This is the density supplied without proof in exercise 2.2, from which we computed 𝔼θ⁢(X(n))=n⁢θ/(n+1)\mathbb{E}_{\theta}(X_{(n)})=n\theta/(n+1) there. The support of the maximum moves with θ\theta, as the support of a single observation does, and exercise 8.2 shows that the maximum is sufficient for θ\theta, applying the factorisation theorem with some care about the support.

The minimum plays the corresponding role for models where the density is largest at the left-hand end of the support. In exercise 8.5 we see that the minimum of an exponential sample is itself exponentially distributed, but with rate multiplied by nn.

Summary.
  • •
    ​

    A statistic TT is sufficient for θ\theta, if the conditional distribution of the data given TT does not depend on θ\theta: the data can be reduced to TT without losing information about θ\theta.

  • •
    ​

    The factorisation theorem states that TT is sufficient if and only if f⁢(x;θ)=g⁢(T⁢(x);θ)⁢h⁢(x)f(x;\theta)=g(T(x);\theta)\,h(x). Using this result, we can see that a unique MLE is always a function of any sufficient statistic.

  • •
    ​

    For exponential families, the summed natural statistic ∑iT⁢(Xi)\sum_{i}T(X_{i}) is sufficient. This includes the Bernoulli total, the Poisson sum and the normal pair (∑iXi,∑iXi2)(\sum_{i}X_{i},\sum_{i}X_{i}^{2}).

  • •
    ​

    A minimal sufficient statistic is a function of any sufficient statistic, and satisfies the condition T⁢(x)=T⁢(y)T(x)=T(y) if and only if the likelihood ratio f⁢(x;θ)/f⁢(y;θ)f(x;\theta)/f(y;\theta) does not depend on θ\theta. A complete statistic is any statistic such that the only function of the statistic which has expectation zero for all θ\theta is the zero function. In lecture 10 we will use this property to identify the unique best unbiased estimator.

  • •
    ​

    The densities of the sample maximum and minimum are n⁢Fn−1⁢fnF^{n-1}f and n⁢(1−F)n−1⁢fn(1-F)^{n-1}f, obtained by differentiating FnF^{n} and 1−(1−F)n1-(1-F)^{n}, respectively.

Exercise 8.1.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Bernoulli with success probability θ\theta and let T=∑i=1nXiT=\sum_{i=1}^{n}X_{i} be the number of successes. From the remark about sums in appendix A we know that T∼Binomial⁢(n,θ)T\sim\text{Binomial}(n,\theta).

  1. 1.
    ​

    Let t∈{0,1,…,n}t\in\{0,1,\dots,n\} and let x∈{0,1}nx\in\{0,1\}^{n} be a data vector with ∑ixi=t\sum_{i}x_{i}=t. Show that ℙθ⁢(X=x|T=t)=1/(nt)\mathbb{P}_{\theta}(X=x\mskip 1.0mu|\mskip 1.0muT=t)=1/\binom{n}{t}.

  2. 2.
    ​

    What is the conditional probability, if ∑ixi≠t\sum_{i}x_{i}\neq t? Using definition 8.1, show that TT is sufficient for θ\theta.

  3. 3.
    ​

    Using the factorisation theorem 8.2, show that the result in the previous part is correct.

Exercise 8.2.

Let X1,…,Xn∼Uniform⁢(0,θ)X_{1},\dots,X_{n}\sim\text{Uniform}(0,\theta) be i.i.d. with θ>0\theta>0.

  1. 1.
    ​

    Using the indicator notation 𝟏{0≤x≤θ}\mathbf{1}_{\{0\leq x\leq\theta\}} for the density of a single observation, show that the joint density can be written as

    f⁢(x;θ)=θ−n⁢ 1{maxi⁡xi≤θ}⁢ 1{mini⁡xi≥0}.f(x;\theta)=\theta^{-n}\,\mathbf{1}_{\{\max_{i}x_{i}\leq\theta\}}\,\mathbf{1}_% {\{\min_{i}x_{i}\geq 0\}}.
  2. 2.
    ​

    Deduce from theorem 8.2 that the sample maximum X(n)=maxi⁡XiX_{(n)}=\max_{i}X_{i} is sufficient for θ\theta. Identify the functions gg and hh.

  3. 3.
    ​

    Why can the factor with mini⁡xi\min_{i}x_{i} be included in hh, but the factor with maxi⁡xi\max_{i}x_{i} cannot?

Exercise 8.3.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. exponential with rate θ>0\theta>0, i.e. f⁢(x;θ)=θ⁢e−θ⁢xf(x;\theta)=\theta e^{-\theta x} for x≥0x\geq 0.

  1. 1.
    ​

    Show that T=∑iXiT=\sum_{i}X_{i} is sufficient for θ\theta.

  2. 2.
    ​

    Using theorem 8.6, show that TT is minimal sufficient.

  3. 3.
    ​

    In example 5.2 we have found the MLE θ^=1/X¯\hat{\theta}=1/\bar{X}. Write θ^\hat{\theta} as a function of TT and explain why the factorisation in the first part of the question guarantees that such a function exists.

Exercise 8.4.

Let X1,…,Xn∼N⁢(μ,σ2)X_{1},\dots,X_{n}\sim N(\mu,\sigma^{2}) be i.i.d.

  1. 1.
    ​

    Assume that μ\mu is known and σ2\sigma^{2} is unknown. Show that ∑i(Xi−μ)2\sum_{i}(X_{i}-\mu)^{2} is sufficient for σ2\sigma^{2}.

  2. 2.
    ​

    Assume that both parameters are unknown. By expanding the squares in the exponent of the joint density, show that the joint density depends on the data only via the pair T=(∑iXi,∑iXi2)T=\bigl{(}\sum_{i}X_{i},\sum_{i}X_{i}^{2}\bigr{)}. Using theorem 8.2, show that TT is sufficient for (μ,σ2)(\mu,\sigma^{2}).

  3. 3.
    ​

    Using theorem 8.6, show that the pair TT from the previous part is minimal sufficient for (μ,σ2)(\mu,\sigma^{2}).

  4. 4.
    ​

    Is the sample mean X¯\bar{X} on its own sufficient for (μ,σ2)(\mu,\sigma^{2}) when both parameters are unknown? Justify your answer using the previous part.

Exercise 8.5.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. exponential with rate θ>0\theta>0, and let X(1)=mini⁡XiX_{(1)}=\min_{i}X_{i}.

  1. 1.
    ​

    Show that X(1)X_{(1)} is exponentially distributed with rate n⁢θn\theta, first by proposition 8.9 and then directly from ℙθ⁢(X(1)>y)\mathbb{P}_{\theta}(X_{(1)}>y).

  2. 2.
    ​

    Show that n⁢X(1)nX_{(1)} is an unbiased estimator for the mean 1/θ1/\theta, and compute its variance.

  3. 3.
    ​

    Compare the variance of n⁢X(1)nX_{(1)} with that of the sample mean X¯\bar{X}, which is also unbiased for 1/θ1/\theta, and comment on the behaviour of the two estimators as nn grows. Is n⁢X(1)nX_{(1)} a function of the sufficient statistic ∑iXi\sum_{i}X_{i}?

Exercise 8.6.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Bernoulli with success probability θ∈(0,1)\theta\in(0,1), where n≥2n\geq 2, and let T=(X1,…,Xn)T=(X_{1},\dots,X_{n}) be the whole sample.

  1. 1.
    ​

    Show that g⁢(T)=X1−X2g(T)=X_{1}-X_{2} satisfies 𝔼θ⁢(g⁢(T))=0\mathbb{E}_{\theta}\bigl{(}g(T)\bigr{)}=0 for all θ\theta and ℙθ⁢(g⁢(T)=0)<1\mathbb{P}_{\theta}\bigl{(}g(T)=0\bigr{)}<1.

  2. 2.
    ​

    Using definition 8.8, show that TT is not complete, despite TT being sufficient.