Lecture 2 Bias, Mean Squared Error and Consistency

We finished the previous lecture with the idea that a good estimator is one which has a sampling distribution concentrated around the true value θ0\theta_{0}, for all possible true values. In this lecture we quantify this idea: the bias measures whether the estimator is on average right, the mean squared error measures how far away the estimator typically is from the truth, and consistency is the property that the estimator converges to the truth as the amount of data increases. The main idea to take away is the decomposition of the mean squared error into a variance and a squared bias.

2.1 Bias and Unbiasedness

The most basic question we can ask about an estimator is whether it gives the correct answer on average.

Definition 2.1.

The bias of an estimator θ^\hat{\theta} for θ\theta is given by

bias(θ^)=𝔼θ⁢(θ^)−θ.\mathop{\mathrm{bias}}\nolimits(\hat{\theta})=\mathbb{E}_{\theta}(\hat{\theta}% )-\theta.

The estimator is unbiased, if θ^\hat{\theta} has finite mean and bias(θ^)=0\mathop{\mathrm{bias}}\nolimits(\hat{\theta})=0 for all θ∈Θ\theta\in\Theta, i.e. if 𝔼θ⁢(θ^)=θ\mathbb{E}_{\theta}(\hat{\theta})=\theta for all θ\theta.

The bias in general depends on the true value θ\theta, as shown by the expectation with θ\theta as a subscript.

The workhorse example is the sample mean. Since we will use the first two moments frequently, we state them here for reference.

Lemma 2.2.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with 𝔼⁢(Xi)=μ\mathbb{E}(X_{i})=\mu and Var(Xi)=σ2\mathop{\mathrm{Var}}\nolimits(X_{i})=\sigma^{2}. Define X¯=(X1+⋯+Xn)/n\bar{X}=(X_{1}+\dots+X_{n})/n. Then

𝔼⁢(X¯)=μandVar(X¯)=σ2n.\mathbb{E}(\bar{X})=\mu\qquad\text{and}\qquad\mathop{\mathrm{Var}}\nolimits(% \bar{X})=\frac{\sigma^{2}}{n}.
Proof.

Using the linearity of expectation we find

𝔼⁢(X¯)=1n⁢∑i=1n𝔼⁢(Xi)=1n⋅n⁢μ=μ.\mathbb{E}(\bar{X})=\frac{1}{n}\sum_{i=1}^{n}\mathbb{E}(X_{i})=\frac{1}{n}% \cdot n\mu=\mu.

Since the XiX_{i} are independent, the variance of the sum can be found as the sum of the variances:

Var(X¯)=1n2⁢∑i=1nVar(Xi)=1n2⋅n⁢σ2=σ2n.\mathop{\mathrm{Var}}\nolimits(\bar{X})=\frac{1}{n^{2}}\sum_{i=1}^{n}\mathop{% \mathrm{Var}}\nolimits(X_{i})=\frac{1}{n^{2}}\cdot n\sigma^{2}=\frac{\sigma^{2% }}{n}.

This completes the proof. ∎

Whenever the parameter of interest is the mean of the observations, as for the Bernoulli and Poisson models from table 1.1, lemma 2.2 shows that 𝔼θ⁢(X¯)=θ\mathbb{E}_{\theta}(\bar{X})=\theta and thus the sample mean is unbiased.

A more delicate example, which we will consider again in section 2.3, is the problem of estimating a variance. If a sample is observed with unknown mean, then one can estimate the variance σ2\sigma^{2} using the sample variance

S2=1n−1⁢∑i=1n(Xi−X¯)2.S^{2}=\frac{1}{n-1}\sum_{i=1}^{n}(X_{i}-\bar{X})^{2}.

The unexpected divisor n−1n-1 instead of the seemingly more obvious nn is required to make this estimator unbiased.

Lemma 2.3.

Let n≥2n\geq 2 and let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with Var(Xi)=σ2\mathop{\mathrm{Var}}\nolimits(X_{i})=\sigma^{2}. Then 𝔼⁢(S2)=σ2\mathbb{E}(S^{2})=\sigma^{2} and thus the sample variance is an unbiased estimator for σ2\sigma^{2}.

Proof.

Let μ=𝔼⁢(Xi)\mu=\mathbb{E}(X_{i}). We prove the algebraic identity

equation (2.1) (2.1)
∑i=1n(Xi−X¯)2=∑i=1n(Xi−μ)2−n⁢(X¯−μ)2,\sum_{i=1}^{n}(X_{i}-\bar{X})^{2}=\sum_{i=1}^{n}(X_{i}-\mu)^{2}-n(\bar{X}-\mu)% ^{2},

which holds for all numbers. Writing Xi−X¯=(Xi−μ)−(X¯−μ)X_{i}-\bar{X}=(X_{i}-\mu)-(\bar{X}-\mu), we can expand the square and use the relation ∑i(Xi−μ)=n⁢(X¯−μ)\sum_{i}(X_{i}-\mu)=n(\bar{X}-\mu) to simplify the expression of the cross term. Taking expectations on (2.1) and using the relations 𝔼⁢((Xi−μ)2)=σ2\mathbb{E}\bigl{(}(X_{i}-\mu)^{2}\bigr{)}=\sigma^{2} as well as 𝔼⁢((X¯−μ)2)=Var(X¯)=σ2/n\mathbb{E}\bigl{(}(\bar{X}-\mu)^{2}\bigr{)}=\mathop{\mathrm{Var}}\nolimits(% \bar{X})=\sigma^{2}/n from lemma 2.2 we find

𝔼⁢(∑i=1n(Xi−X¯)2)=n⁢σ2−n⋅σ2n=(n−1)⁢σ2.\mathbb{E}\Bigl{(}\sum_{i=1}^{n}(X_{i}-\bar{X})^{2}\Bigr{)}=n\sigma^{2}-n\cdot% \frac{\sigma^{2}}{n}=(n-1)\sigma^{2}.

Dividing by n−1n-1 gives the result 𝔼⁢(S2)=σ2\mathbb{E}(S^{2})=\sigma^{2}, as claimed. ∎

The calculation also shows what is wrong with the divisor nn: the estimator (1/n)⁢∑i(Xi−X¯)2(1/n)\sum_{i}(X_{i}-\bar{X})^{2} has expectation (n−1)⁢σ2/n(n-1)\sigma^{2}/n and thus systematically underestimates the variance, since the observations are on average slightly closer to the sample mean X¯\bar{X} than the true mean μ\mu would be. We will return to this comparison, and to the surprising fact that the biased estimator is in some sense better, in section 2.3.

The bias of this estimator is small if the sample is large. A property which is weaker than unbiasedness, but which holds for large sample sizes, is that the bias goes to zero in the limit.

Definition 2.4.

Let (θ^n)n∈ℕ(\hat{\theta}_{n})_{n\in\mathbb{N}} be a family of estimators for θ\theta, depending on the sample size nn. The estimator θ^n\hat{\theta}_{n} is asymptotically unbiased, if

limn→∞bias(θ^n)=0\lim_{n\to\infty}\mathop{\mathrm{bias}}\nolimits(\hat{\theta}_{n})=0

for all θ∈Θ\theta\in\Theta.

Example 2.5.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with Var(Xi)=σ2\mathop{\mathrm{Var}}\nolimits(X_{i})=\sigma^{2} and consider the family of estimators

σ^n2=1n⁢∑i=1n(Xi−X¯)2\hat{\sigma}^{2}_{n}=\frac{1}{n}\sum_{i=1}^{n}(X_{i}-\bar{X})^{2}

for n≥2n\geq 2. Since σ^n2=(n−1)⁢S2/n\hat{\sigma}^{2}_{n}=(n-1)S^{2}/n, we can use lemma 2.3 to conclude that 𝔼⁢(σ^n2)=(n−1)⁢σ2/n\mathbb{E}(\hat{\sigma}^{2}_{n})=(n-1)\sigma^{2}/n and thus

bias(σ^n2)=n−1n⁢σ2−σ2=−σ2n.\mathop{\mathrm{bias}}\nolimits(\hat{\sigma}^{2}_{n})=\frac{n-1}{n}\sigma^{2}-% \sigma^{2}=-\frac{\sigma^{2}}{n}.

While the bias for every nn is non-zero, it converges to 0 as n→∞n\to\infty. Thus, σ^n2\hat{\sigma}^{2}_{n} is a biased but asymptotically unbiased estimator.

2.2 Mean Squared Error

If an estimator is unbiased, this does not mean that it is accurate. The mean squared error is a numerical measure of the typical size of the error θ^−θ\hat{\theta}-\theta.

Definition 2.6.

The mean squared error (MSE) of an estimator θ^\hat{\theta} for θ\theta is given by

MSE(θ^)=𝔼θ⁢((θ^−θ)2).\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})=\mathbb{E}_{\theta}\Bigl{(}\bigl{% (}\hat{\theta}-\theta\bigr{)}^{2}\Bigr{)}.

We want the mean squared error to be small. It combines two different kinds of defects, a systematic offset and random scatter, and the following decomposition separates the two.

Theorem 2.7.

For any estimator θ^\hat{\theta} for θ\theta with finite variance,

MSE(θ^)=Var(θ^)+bias(θ^)2.\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})=\mathop{\mathrm{Var}}\nolimits(% \hat{\theta})+\mathop{\mathrm{bias}}\nolimits(\hat{\theta})^{2}.
Proof.

Let m=𝔼θ⁢(θ^)m=\mathbb{E}_{\theta}(\hat{\theta}). Then we can write bias(θ^)=m−θ\mathop{\mathrm{bias}}\nolimits(\hat{\theta})=m-\theta. Adding and subtracting mm inside the square we get

MSE(θ^)\displaystyle\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})
=𝔼θ⁢(((θ^−m)+(m−θ))2)\displaystyle=\mathbb{E}_{\theta}\Bigl{(}\bigl{(}(\hat{\theta}-m)+(m-\theta)% \bigr{)}^{2}\Bigr{)}
=𝔼θ⁢((θ^−m)2)+2⁢(m−θ)⁢𝔼θ⁢(θ^−m)+(m−θ)2.\displaystyle=\mathbb{E}_{\theta}\bigl{(}(\hat{\theta}-m)^{2}\bigr{)}+2(m-% \theta)\,\mathbb{E}_{\theta}(\hat{\theta}-m)+(m-\theta)^{2}.

Since 𝔼θ⁢(θ^−m)=m−m=0\mathbb{E}_{\theta}(\hat{\theta}-m)=m-m=0, the middle term disappears. The first term is Var(θ^)\mathop{\mathrm{Var}}\nolimits(\hat{\theta}) and the last term is bias(θ^)2\mathop{\mathrm{bias}}\nolimits(\hat{\theta})^{2}. This is the required identity and the proof is complete. ∎

The decomposition allows us to understand the rôle of unbiasedness: for an unbiased estimator, the bias term disappears and we have MSE(θ^)=Var(θ^)\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})=\mathop{\mathrm{Var}}\nolimits(% \hat{\theta}), and thus, when comparing two unbiased estimators, we can only consider the variances. This is the basis of all optimality theory: among the unbiased estimators, the best one is the one with the smallest variance. In lecture 13 we will see that there is a lower bound for the variance of an unbiased estimator.

Using theorem 2.7 and lemma 2.2, we find that the mean squared error of the sample mean is given by MSE(X¯)=σ2/n\mathop{\mathrm{MSE}}\nolimits(\bar{X})=\sigma^{2}/n. For Bernoulli samples this gives MSE(X¯)=θ⁢(1−θ)/n\mathop{\mathrm{MSE}}\nolimits(\bar{X})=\theta(1-\theta)/n and for Poisson samples MSE(X¯)=θ/n\mathop{\mathrm{MSE}}\nolimits(\bar{X})=\theta/n. In both cases the mean squared error decreases as 1/n1/n: more data give a proportionally smaller error.

2.3 When Unbiasedness Is Not Enough

Unbiasedness is an appealing property, but it is a constraint, not a goal, and insisting on it can cost us accuracy. We will show that for the variance of a normal sample the unbiased sample variance is beaten, in terms of mean squared error, by a biased competitor.

Let X1,…,Xn∼N⁢(μ,σ2)X_{1},\dots,X_{n}\sim N(\mu,\sigma^{2}) be i.i.d. and consider the family of estimators

σ^c2=c⁢∑i=1n(Xi−X¯)2,\hat{\sigma}^{2}_{c}=c\sum_{i=1}^{n}(X_{i}-\bar{X})^{2},

for every constant c>0c>0. For c=1/(n−1)c=1/(n-1) this is the unbiased sample variance S2S^{2} and for c=1/nc=1/n it is the underestimating version from the previous section. To compare these estimators, we need to know the distribution of the sum of squares. From proposition A.3 in appendix A we know that the scaled sum of squares satisfies

1σ2⁢∑i=1n(Xi−X¯)2∼χn−12.\frac{1}{\sigma^{2}}\sum_{i=1}^{n}(X_{i}-\bar{X})^{2}\sim\chi^{2}_{n-1}.

If we write W=∑i(Xi−X¯)2W=\sum_{i}(X_{i}-\bar{X})^{2} and recall from definition A.2 that χn−12\chi^{2}_{n-1} has mean n−1n-1 and variance 2⁢(n−1)2(n-1), we get

𝔼⁢(W)=(n−1)⁢σ2andVar(W)=2⁢(n−1)⁢σ4.\mathbb{E}(W)=(n-1)\sigma^{2}\qquad\text{and}\qquad\mathop{\mathrm{Var}}% \nolimits(W)=2(n-1)\sigma^{4}.

Thus, the estimator σ^c2=c⁢W\hat{\sigma}^{2}_{c}=cW has bias bias(σ^c2)=(c⁢(n−1)−1)⁢σ2\mathop{\mathrm{bias}}\nolimits(\hat{\sigma}^{2}_{c})=\bigl{(}c(n-1)-1\bigr{)}% \sigma^{2} and, using theorem 2.7, we get

equation (2.2) (2.2)
MSE(σ^c2)=c2⁢Var(W)+bias(σ^c2)2=(2⁢(n−1)⁢c2+(c⁢(n−1)−1)2)⁢σ4.\mathop{\mathrm{MSE}}\nolimits(\hat{\sigma}^{2}_{c})=c^{2}\,\mathop{\mathrm{% Var}}\nolimits(W)+\mathop{\mathrm{bias}}\nolimits(\hat{\sigma}^{2}_{c})^{2}=% \Bigl{(}2(n-1)c^{2}+\bigl{(}c(n-1)-1\bigr{)}^{2}\Bigr{)}\sigma^{4}.

If we choose c=1/(n−1)c=1/(n-1), the bias term disappears and we get

MSE(S2)=2⁢σ4n−1.\mathop{\mathrm{MSE}}\nolimits(S^{2})=\frac{2\sigma^{4}}{n-1}.

If we choose c=1/nc=1/n, then simplification of (2.2) gives

MSE(σ^1/n2)=2⁢n−1n2⁢σ4.\mathop{\mathrm{MSE}}\nolimits(\hat{\sigma}^{2}_{1/n})=\frac{2n-1}{n^{2}}\,% \sigma^{4}.

The biased estimator wins, since (2⁢n−1)/n2<2/(n−1)(2n-1)/n^{2}<2/(n-1). A little algebra shows that this inequality is equivalent to (2⁢n−1)⁢(n−1)<2⁢n2(2n-1)(n-1)<2n^{2}, i.e. to 2⁢n2−3⁢n+1<2⁢n22n^{2}-3n+1<2n^{2}, which holds for all n≥2n\geq 2.

Remark.

The value c=1/nc=1/n is not optimal. The mean squared error (2.2) is a quadratic function of cc and minimising this function over all c>0c>0 gives c=1/(n+1)c=1/(n+1). For this value we get

MSE(σ^1/(n+1)2)=2⁢σ4n+1.\mathop{\mathrm{MSE}}\nolimits(\hat{\sigma}^{2}_{1/(n+1)})=\frac{2\sigma^{4}}{% n+1}.

Thus, neither the unbiased estimator nor the maximum-likelihood estimator c=1/nc=1/n (which we will see in lecture 5) is optimal, but a third, different estimator exists.

By allowing a small, controlled bias, we have been able to reduce the variance and thus have strictly improved the mean squared error. The lesson to be learned here is not that bias is good, but that an estimator should be evaluated using its mean squared error instead of just the bias.

2.4 Consistency

The mean squared errors we have computed all shrink as the sample grows. Consistency is the property which describes this limiting behaviour: as the number of observations increases, a consistent estimator converges to the true value.

Definition 2.8.

Consider a family (θ^n)n∈ℕ(\hat{\theta}_{n})_{n\in\mathbb{N}} of estimators for θ\theta, one for each sample size nn. The estimator θ^n\hat{\theta}_{n} is said to be consistent for θ\theta, if θ^n→pθ\hat{\theta}_{n}\xrightarrow{\ \mathrm{p}\ }\theta as n→∞n\to\infty, i.e. if for every ε>0\varepsilon>0 we have

limn→∞ℙθ⁢(|θ^n−θ|>ε)=0,\lim_{n\to\infty}\mathbb{P}_{\theta}\bigl{(}|\hat{\theta}_{n}-\theta|>% \varepsilon\bigr{)}=0,

for all θ∈Θ\theta\in\Theta.

Here we write →p\xrightarrow{\ \mathrm{p}\ } for convergence in probability, as recalled in definition A.12 in appendix A: the probability of the estimate being off the truth by any fixed amount eventually becomes negligible. The most basic example of a consistent estimator is the sample mean, where consistency is a manifestation of the law of large numbers.

Theorem 2.9.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with 𝔼θ⁢(Xi)=θ\mathbb{E}_{\theta}(X_{i})=\theta and finite variance σ2\sigma^{2}. Then the sample mean X¯\bar{X} is a consistent estimator for θ\theta.

Proof.

Using Chebyshev’s inequality, lemma A.6 and lemma 2.2, for any ε>0\varepsilon>0 we have

ℙθ⁢(|X¯−θ|>ε)≤Var(X¯)ε2=σ2ε2⁢n⟶0(n→∞).\mathbb{P}_{\theta}\bigl{(}|\bar{X}-\theta|>\varepsilon\bigr{)}\leq\frac{% \mathop{\mathrm{Var}}\nolimits(\bar{X})}{\varepsilon^{2}}=\frac{\sigma^{2}}{% \varepsilon^{2}n}\longrightarrow 0\qquad(n\to\infty).

Thus X¯→pθ\bar{X}\xrightarrow{\ \mathrm{p}\ }\theta and the estimator is consistent. ∎

Remark.

The theorem requires the variance to be finite, as used in the Chebyshev argument; it is sufficient to have the mean finite, by the law of large numbers, theorem A.14 in appendix A. Some condition like this is needed for the result to hold. The standard counterexample is the Cauchy distribution, with density 1/(π⁢(1+x2))1/(\pi(1+x^{2})) on ℝ\mathbb{R} (see table A.2). This distribution has heavy tails and thus 𝔼⁢(|X1|)\mathbb{E}(|X_{1}|) is infinite, i.e. the distribution has no mean. Similarly, the sample mean of a Cauchy distribution is itself Cauchy-distributed, with the same distribution as a single observation for every nn (without proof). Thus, the sample mean does not concentrate as the sample size increases and hence is not a consistent estimator for the centre of the distribution. Consistency of an estimator is a property of the estimator and the model combined, and does not only depend on the form of the formula.

The argument used almost nothing about the sample mean except that its mean squared error converges to zero, and thus the same argument applies to any estimator with this property.

Theorem 2.10.

If MSE(θ^n)→0\mathop{\mathrm{MSE}}\nolimits(\hat{\theta}_{n})\to 0 as n→∞n\to\infty, then θ^n\hat{\theta}_{n} is consistent.

Proof.

Using Markov’s inequality, lemma A.5, for the non-negative random variable (θ^n−θ)2(\hat{\theta}_{n}-\theta)^{2}, for any ε>0\varepsilon>0 we get

ℙθ⁢(|θ^n−θ|>ε)=ℙθ⁢((θ^n−θ)2>ε2)≤𝔼θ⁢((θ^n−θ)2)ε2=MSE(θ^n)ε2⟶0\mathbb{P}_{\theta}\bigl{(}|\hat{\theta}_{n}-\theta|>\varepsilon\bigr{)}=% \mathbb{P}_{\theta}\bigl{(}(\hat{\theta}_{n}-\theta)^{2}>\varepsilon^{2}\bigr{% )}\leq\frac{\mathbb{E}_{\theta}\bigl{(}(\hat{\theta}_{n}-\theta)^{2}\bigr{)}}{% \varepsilon^{2}}=\frac{\mathop{\mathrm{MSE}}\nolimits(\hat{\theta}_{n})}{% \varepsilon^{2}}\longrightarrow 0

as n→∞n\to\infty. Thus θ^n→pθ\hat{\theta}_{n}\xrightarrow{\ \mathrm{p}\ }\theta as required. ∎

Combining this criterion with the bias-variance decomposition gives a sufficient condition in terms of the two quantities we know how to compute.

Corollary 2.11.

If θ^n\hat{\theta}_{n} is asymptotically unbiased and Var(θ^n)→0\mathop{\mathrm{Var}}\nolimits(\hat{\theta}_{n})\to 0 as n→∞n\to\infty, then θ^n\hat{\theta}_{n} is consistent.

Proof.

From theorem 2.7 we know that MSE(θ^n)=Var(θ^n)+bias(θ^n)2\mathop{\mathrm{MSE}}\nolimits(\hat{\theta}_{n})=\mathop{\mathrm{Var}}% \nolimits(\hat{\theta}_{n})+\mathop{\mathrm{bias}}\nolimits(\hat{\theta}_{n})^% {2} and since both terms go to zero by assumption, we can conclude from theorem 2.10 that the result holds. ∎

All estimators in this lecture where we have computed a mean squared error have MSE(θ^n)\mathop{\mathrm{MSE}}\nolimits(\hat{\theta}_{n}) of order 1/n1/n and thus are consistent. Consistency is a coarse property: two consistent estimators can have very different rates of convergence and the mean squared error is a more detailed measure. We will see numerical evidence of consistency in the R session of lecture 9, and in lecture 16 we will see a proof that the maximum likelihood estimator is consistent for a wide class of cases.

Summary.
  • •
    ​

    The bias bias(θ^)=𝔼θ⁢(θ^)−θ\mathop{\mathrm{bias}}\nolimits(\hat{\theta})=\mathbb{E}_{\theta}(\hat{\theta}% )-\theta measures whether an estimator is right on average; an unbiased estimator has zero bias for all θ\theta.

  • •
    ​

    The mean squared error MSE(θ^)=𝔼θ⁢((θ^−θ)2)\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})=\mathbb{E}_{\theta}((\hat{\theta}% -\theta)^{2}) is a measure for the typical squared error, given by MSE(θ^)=Var(θ^)+bias(θ^)2\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})=\mathop{\mathrm{Var}}\nolimits(% \hat{\theta})+\mathop{\mathrm{bias}}\nolimits(\hat{\theta})^{2}.

  • •
    ​

    Unbiasedness is a constraint, not an aim: a biased estimator can have smaller mean squared error, as shown by the normal variance.

  • •
    ​

    An estimator is consistent, if θ^n\hat{\theta}_{n} converges in probability to θ\theta. A sufficient condition for this is MSE(θ^n)→0\mathop{\mathrm{MSE}}\nolimits(\hat{\theta}_{n})\to 0, i.e. being asymptotically unbiased with vanishing variance.

Exercise 2.1.

Let X1,…,XnX_{1},\dots,X_{n} be an i.i.d. sample from the Poisson distribution with parameter θ\theta, i.e. 𝔼θ⁢(Xi)=Varθ(Xi)=θ\mathbb{E}_{\theta}(X_{i})=\mathop{\mathrm{Var}}\nolimits_{\theta}(X_{i})=\theta and let θ^=X¯\hat{\theta}=\bar{X}.

  1. 1.
    ​

    Show that θ^\hat{\theta} is an unbiased estimator for θ\theta.

  2. 2.
    ​

    Determine the mean squared error of θ^\hat{\theta} and show that θ^\hat{\theta} is a consistent estimator.

Exercise 2.2.

Let X1,…,Xn∼Uniform⁢(0,θ)X_{1},\dots,X_{n}\sim\text{Uniform}(0,\theta) be i.i.d., as in example 1.2, and let θ^=maxi⁡Xi\hat{\theta}=\max_{i}X_{i}. You may use that θ^\hat{\theta} has density f⁢(t)=n⁢tn−1/θnf(t)=n\,t^{n-1}/\theta^{n} for 0≤t≤θ0\leq t\leq\theta.

  1. 1.
    ​

    Show that 𝔼θ⁢(θ^)=n⁢θ/(n+1)\mathbb{E}_{\theta}(\hat{\theta})=n\theta/(n+1), and hence that the bias is bias(θ^)=−θ/(n+1)\mathop{\mathrm{bias}}\nolimits(\hat{\theta})=-\theta/(n+1). Is θ^\hat{\theta} asymptotically unbiased?

  2. 2.
    ​

    Show that MSE(θ^)=2⁢θ2/((n+1)⁢(n+2))\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})=2\theta^{2}/\bigl{(}(n+1)(n+2)% \bigr{)}, and deduce that θ^\hat{\theta} is consistent.

  3. 3.
    ​

    Compare the rate at which this mean squared error shrinks with the 1/n1/n rate found for the sample mean, and comment.

Exercise 2.3.

Let X1,…,XnX_{1},\dots,X_{n} be an i.i.d. sample from the exponential distribution with parameter θ\theta, i.e. f⁢(x;θ)=θ⁢e−θ⁢xf(x;\theta)=\theta e^{-\theta x} for x≥0x\geq 0 and 𝔼θ⁢(Xi)=1/θ\mathbb{E}_{\theta}(X_{i})=1/\theta. If we define an estimator by θ^=1/X¯\hat{\theta}=1/\bar{X}, and if we use the fact that T=X1+⋯+XnT=X_{1}+\dots+X_{n} has the gamma density fT⁢(x)=θn⁢xn−1⁢e−θ⁢x/(n−1)!f_{T}(x)=\theta^{n}x^{n-1}e^{-\theta x}/(n-1)! for x≥0x\geq 0, show that, for n≥2n\geq 2,

𝔼θ⁢(θ^)=n⁢θn−1.\mathbb{E}_{\theta}(\hat{\theta})=\frac{n\theta}{n-1}.

From this result, find the bias of θ^\hat{\theta} and show that θ^\hat{\theta} is biased but asymptotically unbiased.