Lecture 2 Bias, Mean Squared Error and Consistency
We finished the previous lecture with the idea that a good estimator is one which has a sampling distribution concentrated around the true value , for all possible true values. In this lecture we quantify this idea: the bias measures whether the estimator is on average right, the mean squared error measures how far away the estimator typically is from the truth, and consistency is the property that the estimator converges to the truth as the amount of data increases. The main idea to take away is the decomposition of the mean squared error into a variance and a squared bias.
2.1 Bias and Unbiasedness
The most basic question we can ask about an estimator is whether it gives the correct answer on average.
The bias of an estimator for is given by
The estimator is unbiased, if has finite mean and for all , i.e. if for all .
The bias in general depends on the true value , as shown by the expectation with as a subscript.
The workhorse example is the sample mean. Since we will use the first two moments frequently, we state them here for reference.
Let be i.i.d. with and . Define . Then
Using the linearity of expectation we find
Since the are independent, the variance of the sum can be found as the sum of the variances:
This completes the proof. ∎
Whenever the parameter of interest is the mean of the observations, as for the Bernoulli and Poisson models from table 1.1, lemma 2.2 shows that and thus the sample mean is unbiased.
A more delicate example, which we will consider again in section 2.3, is the problem of estimating a variance. If a sample is observed with unknown mean, then one can estimate the variance using the sample variance
The unexpected divisor instead of the seemingly more obvious is required to make this estimator unbiased.
Let and let be i.i.d. with . Then and thus the sample variance is an unbiased estimator for .
Let . We prove the algebraic identity
which holds for all numbers. Writing , we can expand the square and use the relation to simplify the expression of the cross term. Taking expectations on (2.1) and using the relations as well as from lemma 2.2 we find
Dividing by gives the result , as claimed. ∎
The calculation also shows what is wrong with the divisor : the estimator has expectation and thus systematically underestimates the variance, since the observations are on average slightly closer to the sample mean than the true mean would be. We will return to this comparison, and to the surprising fact that the biased estimator is in some sense better, in section 2.3.
The bias of this estimator is small if the sample is large. A property which is weaker than unbiasedness, but which holds for large sample sizes, is that the bias goes to zero in the limit.
Let be a family of estimators for , depending on the sample size . The estimator is asymptotically unbiased, if
for all .
Let be i.i.d. with and consider the family of estimators
for . Since , we can use lemma 2.3 to conclude that and thus
While the bias for every is non-zero, it converges to as . Thus, is a biased but asymptotically unbiased estimator.
2.2 Mean Squared Error
If an estimator is unbiased, this does not mean that it is accurate. The mean squared error is a numerical measure of the typical size of the error .
The mean squared error (MSE) of an estimator for is given by
We want the mean squared error to be small. It combines two different kinds of defects, a systematic offset and random scatter, and the following decomposition separates the two.
For any estimator for with finite variance,
Let . Then we can write . Adding and subtracting inside the square we get
Since , the middle term disappears. The first term is and the last term is . This is the required identity and the proof is complete. ∎
The decomposition allows us to understand the rôle of unbiasedness: for an unbiased estimator, the bias term disappears and we have , and thus, when comparing two unbiased estimators, we can only consider the variances. This is the basis of all optimality theory: among the unbiased estimators, the best one is the one with the smallest variance. In lecture 13 we will see that there is a lower bound for the variance of an unbiased estimator.
2.3 When Unbiasedness Is Not Enough
Unbiasedness is an appealing property, but it is a constraint, not a goal, and insisting on it can cost us accuracy. We will show that for the variance of a normal sample the unbiased sample variance is beaten, in terms of mean squared error, by a biased competitor.
Let be i.i.d. and consider the family of estimators
for every constant . For this is the unbiased sample variance and for it is the underestimating version from the previous section. To compare these estimators, we need to know the distribution of the sum of squares. From proposition A.3 in appendix A we know that the scaled sum of squares satisfies
If we write and recall from definition A.2 that has mean and variance , we get
Thus, the estimator has bias and, using theorem 2.7, we get
If we choose , the bias term disappears and we get
If we choose , then simplification of (2.2) gives
The biased estimator wins, since . A little algebra shows that this inequality is equivalent to , i.e. to , which holds for all .
The value is not optimal. The mean squared error (2.2) is a quadratic function of and minimising this function over all gives . For this value we get
Thus, neither the unbiased estimator nor the maximum-likelihood estimator (which we will see in lecture 5) is optimal, but a third, different estimator exists.
By allowing a small, controlled bias, we have been able to reduce the variance and thus have strictly improved the mean squared error. The lesson to be learned here is not that bias is good, but that an estimator should be evaluated using its mean squared error instead of just the bias.
2.4 Consistency
The mean squared errors we have computed all shrink as the sample grows. Consistency is the property which describes this limiting behaviour: as the number of observations increases, a consistent estimator converges to the true value.
Consider a family of estimators for , one for each sample size . The estimator is said to be consistent for , if as , i.e. if for every we have
for all .
Here we write for convergence in probability, as recalled in definition A.12 in appendix A: the probability of the estimate being off the truth by any fixed amount eventually becomes negligible. The most basic example of a consistent estimator is the sample mean, where consistency is a manifestation of the law of large numbers.
Let be i.i.d. with and finite variance . Then the sample mean is a consistent estimator for .
The theorem requires the variance to be finite, as used in the Chebyshev argument; it is sufficient to have the mean finite, by the law of large numbers, theorem A.14 in appendix A. Some condition like this is needed for the result to hold. The standard counterexample is the Cauchy distribution, with density on (see table A.2). This distribution has heavy tails and thus is infinite, i.e. the distribution has no mean. Similarly, the sample mean of a Cauchy distribution is itself Cauchy-distributed, with the same distribution as a single observation for every (without proof). Thus, the sample mean does not concentrate as the sample size increases and hence is not a consistent estimator for the centre of the distribution. Consistency of an estimator is a property of the estimator and the model combined, and does not only depend on the form of the formula.
The argument used almost nothing about the sample mean except that its mean squared error converges to zero, and thus the same argument applies to any estimator with this property.
If as , then is consistent.
Using Markov’s inequality, lemma A.5, for the non-negative random variable , for any we get
as . Thus as required. ∎
Combining this criterion with the bias-variance decomposition gives a sufficient condition in terms of the two quantities we know how to compute.
If is asymptotically unbiased and as , then is consistent.
All estimators in this lecture where we have computed a mean squared error have of order and thus are consistent. Consistency is a coarse property: two consistent estimators can have very different rates of convergence and the mean squared error is a more detailed measure. We will see numerical evidence of consistency in the R session of lecture 9, and in lecture 16 we will see a proof that the maximum likelihood estimator is consistent for a wide class of cases.
-
•
The bias measures whether an estimator is right on average; an unbiased estimator has zero bias for all .
-
•
The mean squared error is a measure for the typical squared error, given by .
-
•
Unbiasedness is a constraint, not an aim: a biased estimator can have smaller mean squared error, as shown by the normal variance.
-
•
An estimator is consistent, if converges in probability to . A sufficient condition for this is , i.e. being asymptotically unbiased with vanishing variance.
Let be an i.i.d. sample from the Poisson distribution with parameter , i.e. and let .
-
1.
Show that is an unbiased estimator for .
-
2.
Determine the mean squared error of and show that is a consistent estimator.
Let be i.i.d., as in example 1.2, and let . You may use that has density for .
-
1.
Show that , and hence that the bias is . Is asymptotically unbiased?
-
2.
Show that , and deduce that is consistent.
-
3.
Compare the rate at which this mean squared error shrinks with the rate found for the sample mean, and comment.
Let be an i.i.d. sample from the exponential distribution with parameter , i.e. for and . If we define an estimator by , and if we use the fact that has the gamma density for , show that, for ,
From this result, find the bias of and show that is biased but asymptotically unbiased.