Lecture 28 Credible Intervals and Bayesian Testing
Lecture 26 ended with the posterior distribution as the complete Bayesian answer. In this lecture we answer the interval question of lecture 25 and the testing question of lectures 20 to 23 in posterior terms, and then set the two approaches side by side on the mean of a normal sample; lecture 27 shows the same comparison in pictures.
We use the two conjugate pairs from lecture 26: the Beta posterior from example 26.5 for the Bernoulli distribution, and for the normal distribution with known variance and prior , we have the posterior from proposition 26.6, with
We consider the data set from example 25.3: observations with and . Assume that before the data were collected we expected the mean to be around and we chose the prior . Then the posterior precision is and the posterior mean is , i.e. the posterior distribution is .
28.1 Credible Intervals
A confidence interval, as given in definition 25.1, is a random interval which covers the fixed parameter with a given probability. In contrast, in a Bayesian model the data are fixed after observation, the parameter is random and an interval estimate is a set which contains most of the posterior probability.
Let be the posterior density of given the data and let . An interval is a credible interval for with credibility , if
The equal-tailed credible interval is the interval , between the - and -quantiles of the posterior distribution.
The probability in the definition is a probability about , for the given data set: the sentence “ is in with probability ” (which we had to forbid in lecture 25) is the exact statement a credible interval makes. The trade-off is that this probability depends on the prior distribution.
Any interval which has posterior probability works, and the equal-tailed interval is the usual choice, because the edges of this interval are quantiles. The alternative would be to find the shortest such set.
A highest posterior density (HPD) set with credibility is the set
where is chosen such that .
The set is the smallest set with posterior probability . Any time we want to exchange part of this set with a region where the density is below , we need to include more length in the interval to make up for the lost probability. For a unimodal posterior distribution, forms an interval, called the HPD interval, where the two boundaries have the same posterior density. For a symmetric, unimodal posterior distribution, this interval will be an equal-tailed interval (see exercise 28.1).
If the posterior distribution is equal to the normal distribution , then the density is symmetric around and decreases as it gets further away from this point. Thus, the HPD interval is the equal-tailed interval
For the data set and prior used in the section, we have the prior and , and the credible interval is , i.e. the interval . The credible interval is shifted towards the prior mean and is shorter than the confidence interval from example 25.3, because the prior provides additional information. In the case of the flat prior, in section 26.4, the posterior distribution is and the credible interval (28.2) coincides with the confidence interval . This coincidence is shown in section 28.3.
For a skewed posterior the HPD interval is shifted towards the mode, slightly shorter than the equal-tailed interval, and the numerical determination of this interval is the topic of exercise 28.2, which shows the difference between the two intervals for the skewed posterior in example 26.5. As increases, the difference between the intervals decreases, because the posterior distribution approaches a symmetric normal distribution and the equal-tailed interval is then appropriate.
28.2 Bayesian Testing
In the Bayesian model a hypothesis , in the sense of definition 20.1, is an event concerning the random variable and thus has a probability before and after the data are seen. Testing does not require new theory.
Let and be hypotheses with positive prior probability. Then the posterior probability of is given by
Similarly, we can find the posterior probability of . A Bayesian test rejects if falls below a threshold or, if both hypotheses are treated equally, chooses the hypothesis with the higher posterior probability.
The threshold plays the role of the level of definition 20.5, but the two numbers are not comparable: bounds the probability of a type I error over repeated samples, whereas the threshold bounds the posterior probability of for the data in hand. Any asymmetry between the hypotheses can be expressed by the prior.
We consider the running data set, testing against . Using the prior , we find the posterior distribution to be . Thus we have
and consequently . A symmetric decision between the hypotheses favours . Using the rule with threshold , we would not have rejected . For the flat prior case, the posterior is and we get . This is a much less certain result than in the example above: the prior , which had only of its mass below , contributed significantly to the evidence against . The same approach can be used to find the posterior probability of a hypothesis about a proportion: for the posterior in example 26.5, the posterior probability of is the value of the Beta distribution function at . This can be computed using any statistical package (in R, the function is pbeta), and the result is . Exercise 28.3 shows how to compute this integral manually.
A simple hypothesis has probability zero under every prior with a density, and thus for two simple hypotheses the prior must be a two-point distribution. The posterior in this case is then simple.
Let and be simple hypotheses and let the prior have probability for and for . Then the prior odds of against are and the posterior odds are .
Using the result of Bayes’ theorem, proposition A.20, we find for and dividing these two expressions we find that the marginal density cancels out. This gives
i.e. the posterior odds are the prior odds times the likelihood ratio. In this context, the likelihood ratio is also called the Bayes factor in favour of . This is the inverse of the ratio considered in the Neyman–Pearson lemma, theorem 22.2. From the posterior odds we can derive the posterior probabilities as .
We consider the running data set again, this time to test the hypothesis against the hypothesis . Since the likelihood of observing a normal sample is from exercise 26.3, we find the Bayes factor to be
Thus we have , and and the exponent is and the Bayes factor is : the data are closer to the value than to and the data favour by a factor of approximately two. If the prior odds are , the posterior odds are and . We see that the Bayes factor multiplies all prior opinions by the same factor, and in cases where the prior is very strong, it can be difficult to overturn the prior opinion. The situation in exercise 28.4 is more complicated, because the prior odds are not equal.
28.3 Comparing Frequentist and Bayesian Answers
Now we have two different sets of answers, for the same data, for the normal mean. We conclude by comparing the two sets.
The first difference is in the nature of the statements. A confidence interval is a statement about a procedure: it says that for repeated samples, the interval covers the fixed truth in of cases. A credible interval is a statement about the parameter, given the data and the prior. Similarly, a -value is computed assuming is true, whereas is the probability of the hypothesis given the data.
The second observation is that for a location parameter with a flat prior, both numbers coincide. Example 28.3 shows the two intervals coinciding, and the tests coincide as well: for the flat-prior posterior probability equals, by the symmetry of , the -value of the one-sided -test from example 20.6; for the running data set both values equal . This is the case because the flat-prior posterior density in has the same shape as the sampling density of in . This is the reason why the misreadings in lectures 20 and 25 were so tempting: the forbidden sentence about the -value is, for the simplest model, a true sentence about the flat-prior posterior, and only the interpretation needs to be changed.
The coincidence is only true for special models and priors. A genuine prior will move the Bayesian answer closer to the prior mean and will concentrate the answer around this value, as in the case of above. If the prior is correct, this is an advantage, but if the prior is wrong this effect will work against us. The frequentist guarantee from definition 25.1 applies for all and thus cannot be violated in this way; exercise 28.5, however, shows that the frequentist coverage of a credible interval can be worse than when is far from the prior mean.
As the sample size increases, the influence of the prior decreases. For the normal mean, equation (28.1) shows that the weight of the prior mean in goes to zero and that so that both the posterior mean and the posterior standard deviation converge to the values they would have with flat prior. The same effect holds for all regular models. The general version of this result is the Bayesian analogue of theorem 17.1: under the regularity conditions from definition 16.1, and for any continuous, positive prior density at the true value, the posterior distribution is approximately
where is the maximum likelihood estimator. Without proof, we state that this result holds. The proof is similar to the one for theorem 17.1: the log-posterior is up to a constant, the log-likelihood is approximately at its maximum, and this quadratic term grows like while does not. From (28.4) we know that the equal-tailed credible interval is approximately the Wald interval of proposition 25.7, independently of the prior: for large sample size both approaches will determine the same interval with approximately frequentist coverage .
The two approaches differ most when testing for a point null hypothesis against a two-sided alternative. In this case a Bayesian test requires a prior which puts a lump of probability on but which spreads the remaining probability over the alternative. The exact distribution of the remaining probability over the alternative matters: if the prior is diffuse on the alternative, can have posterior probability of more than one half for a data set with -value . We will not consider this case in detail, but we can conclude that small -value and small posterior probability of are different quantities and that in this case the difference can be substantial.
-
•
A credible interval with credibility has posterior probability : this is a probability statement about for the given data, taking the prior into account. The equal-tailed interval is based on posterior quantiles, the HPD interval is the shortest one, and for symmetric, unimodal posteriors both intervals coincide.
-
•
For the normal mean with known variance and normal prior, the credible interval is ; for the flat prior this interval coincides with the confidence interval .
-
•
A Bayesian test compares posterior probabilities of the hypotheses, or rejects if falls below a threshold. For two simple hypotheses, the posterior odds are the prior odds times the likelihood ratio, known as the Bayes factor.
-
•
The level of a test and the confidence level of an interval are statements about the procedure applied to repeated samples, whereas the credibility of an interval is a statement about the parameter given one data set. For a location parameter with flat prior, the intervals, and the one-sided test and -value, agree numerically while the statements do not have the same meaning.
-
•
An informative prior will shift and shorten the Bayesian answer; as the posterior is approximately and then the credible and Wald intervals coincide. For point-null hypothesis testing, the two approaches can differ significantly.
Let the posterior density be continuous, strictly increasing on and strictly decreasing on .
-
1.
Show that the HPD set from definition 28.2 is an interval with and .
-
2.
Assume that the density is symmetric about , i.e. that for all . Show that the HPD interval with credibility coincides with the equal-tailed interval .
-
3.
Now assume that the posterior density is continuous and strictly decreasing on and zero elsewhere. Show that the HPD set with credibility is the one-sided interval and that this interval is shorter than the equal-tailed interval.
In example 26.5, where a Bernoulli sample with and successes was combined with the prior to obtain the posterior distribution with mean , the quantiles of the Beta distribution had to be found numerically. Using any statistical package, and using the qbeta function in R, we find the - and -quantiles of to be and .
-
1.
Write down the equal-tailed credible interval for and compare the mid-point of this interval to the posterior mean.
-
2.
Show that the posterior density has its maximum at , i.e. that the posterior distribution is skewed with the mode above the mean.
-
3.
The HPD interval for , found numerically, is . Using definition 28.2, explain why this interval is shifted towards the mode, compared to the equal-tailed interval. Also explain why the difference between the two intervals decreases as increases.
Assume that we have observed two successes in three Bernoulli trials with success probability , and the prior distribution of is flat on the interval .
-
1.
Determine the posterior distribution of .
-
2.
Calculate the posterior probability of by integrating the posterior density. Also, calculate the posterior odds of against .
-
3.
The prior odds of against were . By what factor have the data changed these odds? How is this factor related to the likelihood? Why is it not simply the ratio of the likelihoods at the two parameter values, as in (28.3)?
-
4.
Write down the equations for the boundaries of the equal-tailed credible interval and explain why these equations can only be solved numerically.
In the situation of example 20.6, a sample of size from has mean . We want to compare the simple hypotheses and .
-
1.
Compute the Bayes factor in favour of .
-
2.
Before the data were available, was three times as likely as . Compute the posterior odds and the posterior probability of .
-
3.
In lecture 20 we have seen that the one-sided -value can be found for the data. Compare your two numbers and explain why they can disagree, despite the fact that both values are against .
-
4.
What value must the prior odds in favour of have for the posterior probability of to be greater than one half?
Let be i.i.d. with where is known, and let the prior be . Let be the weight of the prior mean in . Consider the credible interval (28.2) to be a random interval, as introduced in lecture 25, where is fixed and the data are random.
-
1.
Show that and that under the posterior mean satisfies .
-
2.
Determine the coverage probability as a function of .
-
3.
For , , , and , determine the coverage probability for and for and discuss the result as .
-
4.
Is the credible interval a confidence interval with level in the sense of definition 25.1? What happens to the coverage probability as with fixed?