Lecture 27 Confidence Versus Credibility
This is the R session for week 9 of the course. In lectures this week we have introduced the topics of confidence intervals and Bayesian priors and posteriors. The aim of this R session is to visualise these concepts: we will consider a hundred confidence intervals, together with a count of how often these intervals “miss” the target, we will consider a confidence interval and a Bayesian interval for the same data, and we will consider how a prior can be transformed into a posterior as observations are added one by one. Today we will use the Bayesian interval in an informal way, in preparation for the fact that in one week (in lecture 28) we will introduce the concept of a credible interval and will then be able to make a proper comparison between confidence intervals and credible intervals.
27.1 Coverage by Simulation
The interval of the day is the -interval from example 25.3, using the numerical values from the example (, are known), so that the interval is . By definition 25.1 of the level , the random interval contains the fixed value in of all samples. To verify this, we consider and compute the interval for samples. Before you execute the code, answer the following questions: (1) How many of the intervals will not cover the value ? (2) Will all intervals have the same width? (3) Will the misses be clustered or evenly distributed among the 100 samples?
As in lecture 12, covers is a vector of TRUE and FALSE, one for each interval, indicating whether the interval covers the truth. The command mean counts the proportion of TRUE values, which are interpreted as ones and zeros by R. For the example, out of intervals cover the truth. The misses are for samples , , , and . The exact number is a coincidence, determined by the seed of the random number generator: the number of intervals which cover the truth is binomially distributed with standard deviation , so or would have been just as unremarkable. The command segments can be used to plot each interval as a horizontal line, and the samples where the interval covers the truth are plotted in grey while the misses are plotted in black.

Figure 27.1 illustrates the confidence level of the procedure in a graphical way. The intervals in the figure all have the same width, since the value is known: the intervals move, while the dashed line stays fixed. The misses are not clustered but are scattered throughout the figure, since the samples are independent. If we had only observed sample , the interval would have been indistinguishable from the others, and there would have been no indication in the data that this was a miss. For this reason, the statement “ is in the interval with probability ” is wrong: the value is either in the interval or it is not, but in this instance it is not. The probability refers to the procedure, not to the individual values. Before the code is executed, predict the long-run proportion of intervals which cover the truth, and determine the distance from where an error is no longer considered significant.
The long-run coverage is . The standard error of this estimate is , so the estimate is about one and a half standard errors below the nominal value : this is good agreement, considering the simulation error. The pivot argument in example 25.3 guarantees that this result holds for every value of .
27.2 The Same Data, Two Intervals
We now take the data of example 25.8, successes in Bernoulli trials, and compute two intervals for the success probability . The first is the Wald interval of that example. For the second we treat as random, as in lecture 26: the flat-prior posterior is by example 26.5, and the interval between its and quantiles contains with posterior probability ; this is the equal-tailed credible interval of lecture 28, and the command qbeta can be used to find the quantiles. Before you run the code, predict whether the credible interval will be wider or narrower than the Wald interval, and whether it will sit to the left or to the right of it.
The two intervals, and , coincide to two decimal digits; the credible interval is slightly shifted towards , due to the fact that the flat prior adds one success and one failure to the data. The Wald interval states that the recipe used to compute it covers the truth in approximately of samples. It does not make a statement about this particular interval. The credible interval states that, given the observations and the flat prior, the probability that falls between and is : the sentence forbidden in the previous section is the correct interpretation in this context, since is a random variable in the Bayesian model. The posterior distribution can be used to answer other questions; the command pbeta can be used to evaluate its distribution function.
The Wald interval itself has posterior probability under the flat prior, and the posterior probability that exceeds is . The frequentist interval cannot make a statement about this probability. The final line shows the price of departing from the flat prior: under the informative prior , with mean and equivalent sample size , the same data give the credible interval . A credible interval always depends on the prior, and numerical agreement with a confidence interval is a property of the flat prior and of the sample size. We will consider this comparison in more detail in lecture 28.
27.3 From Prior to Posterior
The following picture illustrates the update from lecture 26, where we process the observations one by one. In the example, we simulate Bernoulli observations with . After observations, with successes, the prior distribution with has evolved into the posterior distribution with mean . This result was obtained in example 26.5. The box below contains a script for computing the posterior mean of the flat prior after observations. The lines in the box are shuffled, so you will need to sort them into a working script before you can try the code.
As always, names must be defined before they are used, and the seed must be set before the first random draw: theta first, followed by the seed, then rbinom with size = 1 to generate the Bernoulli draws, then cumsum to generate the running total of successes; a, b and k can be placed anywhere before the last line. The script R/S27-posterior.R extends this to handle three priors simultaneously, stored as the rows of a matrix: the flat prior , the Jeffreys prior from example 26.10 and the misplaced prior from the previous section. Before you look at the output of the following code, predict the result: which of the three posterior means will be closest to after five observations? Will you still be able to distinguish between the three posterior means after one hundred observations?
The start of the sequence is bad: the first five observations are all failures and after twenty observations there have been only two successes. The probability of at most two successes in twenty observations is for . For this case, the prediction at is dominated by the flat prior: the posterior mean is closest to , since five failures have driven both weak priors below the truth, and the Jeffreys prior is driven lower than the flat prior is. The misplaced prior is still at . The two weak priors follow the data and end at and , both far below the truth. The misplaced prior, moving from to , is the closest of the three at : with twenty observations, the data alone is not enough to avoid being misled, but a prior in the correct region helps and a prior in the wrong region does harm. The data does not allow us to infer the correct region. By the two weak priors are indistinguishable at and even the misplaced prior, with an equivalent sample size of , is now mostly aligned with the sample proportion. The script plots the three posterior densities at each stage using curve and dbeta. Before you look at figure 27.2, sketch the panel for : which curves are steep, where do they take their peaks and how wide are the curves?

Figure 27.2 summarises the update procedure. At the three curves correspond to the prior distributions from lecture 26. One can see that the Jeffreys prior is not flat. At the two weak posteriors have the largest values at , since no successes have been observed so far, and the informative posterior has not yet significantly changed. At all three posteriors have the form of humps, and the posteriors are narrower than in the previous step: as data are added, the posterior gets sharper, irrespective of the prior distribution. At the two weak posteriors coincide and the informative posterior overlaps these two: this is a consequence of the prior being washed out in example 26.5, and the panel for serves as a reminder that data are needed to achieve this effect.
Repeat the picture in figure 27.1 for the -interval from example 25.4, but with , sigma replaced with sd(x) and qnorm(0.975) replaced with qt(0.975, df = 4). The widths of the intervals will change. Estimate how often the intervals miss the target value, and estimate the long-run coverage from samples. Repeat the experiment, but this time leaving qnorm(0.975) instead of the quantile, and comment on your result. Finally, choose a prior mean and an equivalent sample size as in exercise 26.1, use your student ID as the seed for the random number generator, and repeat the posterior sequence of the last section, both with and without the prior. For both cases, determine the first where the posterior mean is within of the truth.
Type the commands used in this session into R, and check that you get the same results as the text.
-
•
The written reference for the commands used in this session is appendix B.
-
•
The confidence level of a statistical procedure is a property of the procedure: in simulated trials of intervals, intervals contained the truth, the misses were indistinguishable from the successes, and in the long run the coverage was . A computed confidence interval either contains the parameter or it does not.
-
•
For the same Bernoulli data, the Wald interval and the credible interval with flat prior coincide to two decimal digits, but they are different statements: the credible interval is a probability statement about , conditioned on the data and on the prior, and thus depends on the prior.
-
•
The Beta posterior gets more and more sharp as observations are added, and in the limit it converges to the sample proportion for every prior. For observations the prior is still noticeable, but for the prior is largely washed out.