Lecture 25 Confidence Intervals
So far we have considered estimators which output a single number. In this lecture we will consider intervals of parameter values, constructed such that the true parameter value lies inside the interval with a given probability. We will see two different ways to construct such an interval, one using the concept of a pivotal quantity and the other one based on the inversion of a family of hypothesis tests. This latter approach also allows us to relate the testing theory from lectures 20 to 23 to the topic of estimation. The Bayesian equivalents of confidence intervals are called credible intervals and we will learn about these in lecture 28. In lecture 27 we will compare the two approaches to confidence intervals using simulation.
25.1 Definition and Interpretation
Here, is a sample from a parametric model with parameter , and is the distribution of the sample under . For a fixed small , typically , the aim is to find an interval which misses the truth with probability at most .
Let and be statistics such that . The random interval is a confidence interval for with confidence level , if
The probability on the left-hand side is called the coverage probability of the interval at .
The endpoints of the interval are statistics, i.e. functions of the data, while the parameter is an unknown constant which does not change between samples. For continuous data we can usually achieve equality in the definition, but for discrete data the coverage probability of the interval can jump as changes. For this reason, the definition only requires the left-hand side to satisfy the inequality. More generally, any random subset of the parameter space such that for all is called a confidence set.
Some care is needed when interpreting a confidence interval. For example, if the data lead to the interval , both endpoints are now fixed numbers, and so is the unknown ; the statement is either true or false. The describes the procedure: of the computed intervals would cover the truth in repeated samples, but we do not know whether the interval in front of us is one of these. The statement “ is in with probability ” is wrong, since it assumes that the parameter is random (this only holds in the Bayesian model from lecture 28). The statement “ of the observations are in the interval” is wrong, since this would be a confusion of parameter with data. The correct interpretation of such an interval can be illustrated by simulating many intervals and seeing how many of these cover the truth. This is done in lecture 27.
25.2 Pivotal Quantities
The most direct approach to constructing a confidence interval is to construct an interval based on a quantity whose distribution is known exactly, regardless of the parameter value. Such a quantity is called a pivotal quantity.
A function of the sample and the parameter is a pivotal quantity or pivot, if the distribution of under does not depend on .
A pivot is not a statistic, because it contains the unknown , but we know the quantiles of the pivot distribution. To get a confidence interval, we can choose constants such that and then solve the inequalities for . If is monotone in , the solution of these inequalities forms an interval which, with probability exactly , contains the unknown parameter value. Using the quantile notation from lecture 20, we have and, using symmetry, . Similarly, we write and for the -quantiles of the and the distribution, respectively. These quantiles are all tabulated.
Let be i.i.d. where is known. Then, by proposition A.3, and thus
for all . Thus, is a pivot. We can find a confidence interval for by bracketing between and and then solving the resulting inequalities for . The resulting confidence interval is
with level . For observations, with and , the interval is . This is the same interval as we discussed above.
If is unknown, this pivot is not available. If we replace by the sample standard deviation , we get a new pivot and we have to consider the distribution from appendix A.
Let be i.i.d. , where both parameters are unknown, and let be the sample variance. By proposition A.3, the two values
are independent. By definition A.4, the ratio
is distributed as and, since the unknown cancels, is a pivot for , independently of the nuisance parameter . The -interval
with level , can be obtained by bracketing the pivot as above. For the data from example 25.3, with sample standard deviation , we can use the table to find and the interval . The wider interval is the price we pay for not knowing . As increases, this difference decreases, since converges to the standard normal distribution.
In the same setting, is a pivot for , independently of the unknown . If we bracket this value between and and then solve for , the order of the endpoints is reversed, and we obtain the confidence interval
for with level . For and the table gives and , and thus the interval is given by
This interval is far from symmetric around the estimate .
If we take the square root of both endpoints, the interval for has the same coverage probability: applying a strictly increasing function to both endpoints does not change the coverage probability, whereas applying a decreasing function swaps the endpoints.
25.3 Inverting a Test
A level- test of decides whether a single value is compatible with the data; a confidence interval collects all the values which are. The following theorem, using the notions of critical region and level from definitions 20.2 and 20.5, makes this correspondence exact, in both directions.
For every , let be the acceptance region of a level- test of , that is the set of samples for which is not rejected. Then
is a confidence set for with level . Conversely, if is a confidence set with level , then for each the test which rejects if and only if has level .
By the definition of , for every and every sample we have if and only if . Let be the true value. Then, for , under the null hypothesis is true and the test with acceptance region rejects with probability at most . We have
Since was arbitrary, is a confidence set with level . For the converse statement, the test for rejects if and only if . Thus, under , this event has probability and the test has level . This completes the proof. ∎
The theorem is a change of viewpoint, rather than a new result: the event “ is not rejected by ” can be read both as a test (with fixed and random) and as a confidence set (with fixed and varying). As an illustration, the size- two-sided test from example 20.7 for the mean of an i.i.d. sample with known rejects if and the values not rejected by the test form the interval from example 25.3. Similarly, by inverting the one-sided test from example 20.6, we find a one-sided confidence interval. We will consider this type of interval again in section 25.5. Finally, in exercise 25.4, the same approach is repeated with unknown . In practice, the duality is most often used in the opposite direction: a value is rejected at level if it does not belong to the interval.
25.4 Approximate Intervals from the MLE
Exact pivots only exist for special models. For every regular model, however, corollary 17.3 provides an approximate pivot: the standardised error of the maximum likelihood estimator, with , converges in distribution to whatever the value of .
With the same assumptions as in corollary 17.3, the Wald interval
has coverage probability that converges to as , for all .
Let and . The event that the interval contains is the event . This event is the same as . From corollary 17.3 we know that with . Writing for the distribution function of , for every we have
Since is continuous everywhere, definition A.13 gives for every , and thus the lower bound converges to and the upper bound converges to as . Letting and using the continuity of again, we find that as and thus the coverage probability converges to . This is the required result. ∎
The Wald interval, for , is the interval . This is the interval introduced in lectures 12 and 18. Using the delta method, theorem A.18, the same approach can be used to derive the interval for any smooth function (see exercise 25.3). However, the coverage of this interval is only guaranteed in the limit, and section 17.5 lists the cases to be aware of.
Let be i.i.d. Bernoulli with success probability . The MLE is by example 5.3 and the Fisher information is by example 11.5. Thus the Wald interval is . For successes in trials we have , the standard error is , and the interval is . If is close to or , or if is small, the coverage of this interval may be less than , since the normal approximation to is poor in these cases and the standard error equals zero if all observations are the same.
25.5 One-Sided Intervals and the Choice of Interval
Sometimes it is only important to consider one direction of a confidence interval, for example when we want to show that a failure rate is less than or equal to some value. A lower confidence bound for with level is a statistic such that for all . The corresponding one-sided confidence interval is then . An upper confidence bound is defined symmetrically. In the pivot method, a one-sided interval can be obtained by considering the single inequality , where is the -quantile of the pivot. For example, for the normal distribution with known variance, this gives the lower bound as promised in section 25.3, where is used instead of since all of the error probability is allocated to one side. In the context of the pivot method, any pair of values with defines a valid two-sided interval. Here we use the equal-tailed choice, i.e. and are the - and -quantiles, respectively. This choice is the shortest interval for a symmetric, unimodal pivot density, but may not be shortest for skewed densities like the chi-squared distribution; the interval boundaries can be found in the tables. Exercise 25.2 shows an example where the shortest interval is one-sided. In lecture 28 we will see the same choice of interval again, when we consider credible intervals.
-
•
A confidence interval covers the true parameter with probability at least for all . The interval is random, the parameter is fixed, and the level describes the procedure, not the individual interval.
-
•
An exact interval can be obtained by bracketing the pivot between two quantiles and solving for . For normal samples this can be used to derive the -, - and chi-squared intervals.
-
•
Tests and confidence sets are dual: the values not rejected by a family of level- tests form a confidence set, and conversely.
-
•
The Wald interval asymptotically covers for all regular models, but the finite-sample coverage may be lower.
-
•
One-sided intervals use instead of . For two-sided intervals the equal-tailed choice is the norm.
Let be i.i.d. exponential with rate and let . In exercise 13.1 we have seen that .
-
1.
Explain why is a pivot for and derive the equal-tailed confidence interval for with level .
-
2.
For observations , compute the interval, using the values and .
- 3.
Let be i.i.d. uniformly distributed on the set and let .
-
1.
Using the density of from example 8.10, show that for and conclude that is a pivot.
-
2.
Show that is a confidence interval for with level and compute the interval for , and .
-
3.
Among all intervals with and , show that the interval from the previous part is the shortest. (Hint: think about how the length of the interval changes when is moved and is adjusted to maintain the same coverage.)
Let be i.i.d. Poisson with mean . Then the MLE is and the Fisher information is from example 11.6.
-
1.
Find the Wald interval for with level and determine the interval for level when is the number of counts with .
-
2.
In exercise 17.2 we have found the asymptotic distribution of the MLE for . Use this distribution to find an approximate confidence interval for and work out the interval for the given data.
-
3.
Another interval for can be found by applying the decreasing function to the boundaries of the interval from the first part. Show that this interval also has asymptotic coverage , compute the interval for the given data and comment on the differences between the two intervals.
Let be i.i.d. , where both parameters are unknown. For consider the test which rejects if .
-
1.
Show that the test has size , independent of the value of .
- 2.
-
3.
For the data given in example 25.4, can you decide whether and should be rejected at level , without having to compute the test statistic?
-
4.
Now consider the one-sided test which rejects if . Invert this test and describe the resulting confidence set.