Lecture 28 Credible Intervals and Bayesian Testing

Lecture 26 ended with the posterior distribution as the complete Bayesian answer. In this lecture we answer the interval question of lecture 25 and the testing question of lectures 20 to 23 in posterior terms, and then set the two approaches side by side on the mean of a normal sample; lecture 27 shows the same comparison in pictures.

We use the two conjugate pairs from lecture 26: the Beta posterior from example 26.5 for the Bernoulli distribution, and for the normal distribution X1,…,Xn∼N⁢(θ,σ2)X_{1},\dots,X_{n}\sim N(\theta,\sigma^{2}) with known variance and prior θ∼N⁢(m,τ2)\theta\sim N(m,\tau^{2}), we have the posterior N⁢(mn,τn2)N(m_{n},\tau_{n}^{2}) from proposition 26.6, with

equation (28.1) (28.1)
1τn2=1τ2+nσ2andmn=τn2⁢(mτ2+n⁢x¯σ2).\frac{1}{\tau_{n}^{2}}=\frac{1}{\tau^{2}}+\frac{n}{\sigma^{2}}\qquad\text{and}% \qquad m_{n}=\tau_{n}^{2}\Bigl{(}\frac{m}{\tau^{2}}+\frac{n\bar{x}}{\sigma^{2}% }\Bigr{)}.

We consider the data set from example 25.3: n=16n=16 observations with σ=2\sigma=2 and x¯=10.3\bar{x}=10.3. Assume that before the data were collected we expected the mean to be around 1212 and we chose the prior N⁢(12,1)N(12,1). Then the posterior precision is 1/1+16/4=51/1+16/4=5 and the posterior mean is mn=0.2⋅(12+4⋅10.3)=10.64m_{n}=0.2\cdot(12+4\cdot 10.3)=10.64, i.e. the posterior distribution is N⁢(10.64,0.2)N(10.64,0.2).

28.1 Credible Intervals

A confidence interval, as given in definition 25.1, is a random interval which covers the fixed parameter with a given probability. In contrast, in a Bayesian model the data are fixed after observation, the parameter is random and an interval estimate is a set which contains most of the posterior probability.

Definition 28.1.

Let π⁢(θ|x)\pi(\theta\mskip 1.0mu|\mskip 1.0mux) be the posterior density of θ\theta given the data xx and let α∈(0,1)\alpha\in(0,1). An interval [l,u][l,u] is a credible interval for θ\theta with credibility 1−α1-\alpha, if

ℙ⁢(l≤θ≤u|x)=∫luπ⁢(ϑ|x)⁢dϑ=1−α.\mathbb{P}(l\leq\theta\leq u\mskip 1.0mu|\mskip 1.0mux)=\int_{l}^{u}\pi(% \vartheta\mskip 1.0mu|\mskip 1.0mux)\,\mathrm{d}\vartheta=1-\alpha.

The equal-tailed credible interval is the interval [qα/2,q1−α/2][q_{\alpha/2},q_{1-\alpha/2}], between the α/2\alpha/2- and (1−α/2)(1-\alpha/2)-quantiles of the posterior distribution.

The probability in the definition is a probability about θ\theta, for the given data set: the sentence “θ\theta is in [l,u][l,u] with probability 0.950.95” (which we had to forbid in lecture 25) is the exact statement a 95%95\% credible interval makes. The trade-off is that this probability depends on the prior distribution.

Any interval which has posterior probability 1−α1-\alpha works, and the equal-tailed interval is the usual choice, because the edges of this interval are quantiles. The alternative would be to find the shortest such set.

Definition 28.2.

A highest posterior density (HPD) set with credibility 1−α1-\alpha is the set

Hc={θ∈Θ|π⁢(θ|x)≥c},H_{c}=\bigl{\{}\,\theta\in\Theta\mathrel{\big{|}}\pi(\theta\mskip 1.0mu|\mskip 1% .0mux)\geq c\,\bigr{\}},

where c>0c>0 is chosen such that ℙ⁢(θ∈Hc|x)=1−α\mathbb{P}(\theta\in H_{c}\mskip 1.0mu|\mskip 1.0mux)=1-\alpha.

The set HcH_{c} is the smallest set with posterior probability 1−α1-\alpha. Any time we want to exchange part of this set with a region where the density is below cc, we need to include more length in the interval to make up for the lost probability. For a unimodal posterior distribution, HcH_{c} forms an interval, called the HPD interval, where the two boundaries have the same posterior density. For a symmetric, unimodal posterior distribution, this interval will be an equal-tailed interval (see exercise 28.1).

Example 28.3.

If the posterior distribution is equal to the normal distribution N⁢(mn,τn2)N(m_{n},\tau_{n}^{2}), then the density is symmetric around mnm_{n} and decreases as it gets further away from this point. Thus, the HPD interval is the equal-tailed interval

equation (28.2) (28.2)
[mn−z1−α/2⁢τn,mn+z1−α/2⁢τn].\bigl{[}\,m_{n}-z_{1-\alpha/2}\,\tau_{n},\ m_{n}+z_{1-\alpha/2}\,\tau_{n}\,% \bigr{]}.

For the data set and prior used in the section, we have the prior N⁢(12,1)N(12,1) and τn=0.2=0.447\tau_{n}=\sqrt{0.2}=0.447, and the 95%95\% credible interval is 10.64±1.96⋅0.44710.64\pm 1.96\cdot 0.447, i.e. the interval [9.76,11.52][9.76,11.52]. The credible interval is shifted towards the prior mean and is shorter than the 95%95\% confidence interval [9.32,11.28][9.32,11.28] from example 25.3, because the prior provides additional information. In the case of the flat prior, τ→∞\tau\to\infty in section 26.4, the posterior distribution is N⁢(x¯,σ2/n)N(\bar{x},\sigma^{2}/n) and the credible interval (28.2) coincides with the confidence interval x¯±z1−α/2⁢σ/n\bar{x}\pm z_{1-\alpha/2}\,\sigma/\sqrt{n}. This coincidence is shown in section 28.3.

For a skewed posterior the HPD interval is shifted towards the mode, slightly shorter than the equal-tailed interval, and the numerical determination of this interval is the topic of exercise 28.2, which shows the difference between the two intervals for the skewed posterior Beta⁢(9,5)\text{Beta}(9,5) in example 26.5. As nn increases, the difference between the intervals decreases, because the posterior distribution approaches a symmetric normal distribution and the equal-tailed interval is then appropriate.

28.2 Bayesian Testing

In the Bayesian model a hypothesis H0:θ∈Θ0H_{0}\colon\theta\in\Theta_{0}, in the sense of definition 20.1, is an event concerning the random variable θ\theta and thus has a probability before and after the data are seen. Testing does not require new theory.

Definition 28.4.

Let H0:θ∈Θ0H_{0}\colon\theta\in\Theta_{0} and H1:θ∈Θ1H_{1}\colon\theta\in\Theta_{1} be hypotheses with positive prior probability. Then the posterior probability of H0H_{0} is given by

ℙ⁢(H0|x)=ℙ⁢(θ∈Θ0|x)=∫Θ0π⁢(ϑ|x)⁢dϑ.\mathbb{P}(H_{0}\mskip 1.0mu|\mskip 1.0mux)=\mathbb{P}(\theta\in\Theta_{0}% \mskip 1.0mu|\mskip 1.0mux)=\int_{\Theta_{0}}\pi(\vartheta\mskip 1.0mu|\mskip 1% .0mux)\,\mathrm{d}\vartheta.

Similarly, we can find the posterior probability of H1H_{1}. A Bayesian test rejects H0H_{0} if ℙ⁢(H0|x)\mathbb{P}(H_{0}\mskip 1.0mu|\mskip 1.0mux) falls below a threshold or, if both hypotheses are treated equally, chooses the hypothesis with the higher posterior probability.

The threshold plays the role of the level α\alpha of definition 20.5, but the two numbers are not comparable: α\alpha bounds the probability of a type I error over repeated samples, whereas the threshold bounds the posterior probability of H0H_{0} for the data in hand. Any asymmetry between the hypotheses can be expressed by the prior.

Example 28.5.

We consider the running data set, testing H0:θ≤10H_{0}\colon\theta\leq 10 against H1:θ>10H_{1}\colon\theta>10. Using the prior N⁢(12,1)N(12,1), we find the posterior distribution to be N⁢(10.64,0.2)N(10.64,0.2). Thus we have

ℙ⁢(H0|x)=ℙ⁢(θ≤10|x)=Φ⁢(10−10.640.447)=Φ⁢(−1.43)=0.076,\mathbb{P}(H_{0}\mskip 1.0mu|\mskip 1.0mux)=\mathbb{P}(\theta\leq 10\mskip 1.0% mu|\mskip 1.0mux)=\Phi\Bigl{(}\frac{10-10.64}{0.447}\Bigr{)}=\Phi(-1.43)=0.076,

and consequently ℙ⁢(H1|x)=0.924\mathbb{P}(H_{1}\mskip 1.0mu|\mskip 1.0mux)=0.924. A symmetric decision between the hypotheses favours H1H_{1}. Using the rule with threshold 0.050.05, we would not have rejected H0H_{0}. For the flat prior case, the posterior is N⁢(10.3,0.25)N(10.3,0.25) and we get ℙ⁢(H0|x)=Φ⁢(−0.6)=0.274\mathbb{P}(H_{0}\mskip 1.0mu|\mskip 1.0mux)=\Phi(-0.6)=0.274. This is a much less certain result than in the example above: the prior N⁢(12,1)N(12,1), which had only 2.3%2.3\% of its mass below 1010, contributed significantly to the evidence against H0H_{0}. The same approach can be used to find the posterior probability of a hypothesis about a proportion: for the posterior Beta⁢(9,5)\text{Beta}(9,5) in example 26.5, the posterior probability of H0:θ≤1/2H_{0}\colon\theta\leq 1/2 is the value of the Beta distribution function at 1/21/2. This can be computed using any statistical package (in R, the function is pbeta), and the result is 0.1330.133. Exercise 28.3 shows how to compute this integral manually.

A simple hypothesis θ=θ0\theta=\theta_{0} has probability zero under every prior with a density, and thus for two simple hypotheses the prior must be a two-point distribution. The posterior in this case is then simple.

Definition 28.6.

Let H0:θ=θ0H_{0}\colon\theta=\theta_{0} and H1:θ=θ1H_{1}\colon\theta=\theta_{1} be simple hypotheses and let the prior have probability p0p_{0} for θ0\theta_{0} and p1=1−p0p_{1}=1-p_{0} for θ1\theta_{1}. Then the prior odds of H0H_{0} against H1H_{1} are p0/p1p_{0}/p_{1} and the posterior odds are ℙ⁢(H0|x)/ℙ⁢(H1|x)\mathbb{P}(H_{0}\mskip 1.0mu|\mskip 1.0mux)/\mathbb{P}(H_{1}\mskip 1.0mu|% \mskip 1.0mux).

Using the result of Bayes’ theorem, proposition A.20, we find ℙ⁢(Hj|x)=pj⁢f⁢(x|θj)/m⁢(x)\mathbb{P}(H_{j}\mskip 1.0mu|\mskip 1.0mux)=p_{j}\,f(x\mskip 1.0mu|\mskip 1.0% mu\theta_{j})/m(x) for j=0,1j=0,1 and dividing these two expressions we find that the marginal density m⁢(x)m(x) cancels out. This gives

equation (28.3) (28.3)
ℙ⁢(H0|x)ℙ⁢(H1|x)=p0p1⋅f⁢(x|θ0)f⁢(x|θ1),\frac{\mathbb{P}(H_{0}\mskip 1.0mu|\mskip 1.0mux)}{\mathbb{P}(H_{1}\mskip 1.0% mu|\mskip 1.0mux)}=\frac{p_{0}}{p_{1}}\cdot\frac{f(x\mskip 1.0mu|\mskip 1.0mu% \theta_{0})}{f(x\mskip 1.0mu|\mskip 1.0mu\theta_{1})},

i.e. the posterior odds are the prior odds times the likelihood ratio. In this context, the likelihood ratio is also called the Bayes factor in favour of H0H_{0}. This is the inverse of the ratio considered in the Neyman–Pearson lemma, theorem 22.2. From the posterior odds rr we can derive the posterior probabilities as ℙ⁢(H0|x)=r/(1+r)\mathbb{P}(H_{0}\mskip 1.0mu|\mskip 1.0mux)=r/(1+r).

Example 28.7.

We consider the running data set again, this time to test the hypothesis H0:θ=10H_{0}\colon\theta=10 against the hypothesis H1:θ=11H_{1}\colon\theta=11. Since the likelihood of observing a normal sample is f⁢(x|θ)∝exp⁡(−n⁢(x¯−θ)2/(2⁢σ2))f(x\mskip 1.0mu|\mskip 1.0mu\theta)\propto\exp\bigl{(}-n(\bar{x}-\theta)^{2}/(% 2\sigma^{2})\bigr{)} from exercise 26.3, we find the Bayes factor to be

f⁢(x|θ0)f⁢(x|θ1)=exp⁡(n2⁢σ2⁢((x¯−θ1)2−(x¯−θ0)2))=exp⁡(n⁢(θ1−θ0)⁢(θ0+θ1−2⁢x¯)2⁢σ2).\frac{f(x\mskip 1.0mu|\mskip 1.0mu\theta_{0})}{f(x\mskip 1.0mu|\mskip 1.0mu% \theta_{1})}=\exp\Bigl{(}\frac{n}{2\sigma^{2}}\bigl{(}(\bar{x}-\theta_{1})^{2}% -(\bar{x}-\theta_{0})^{2}\bigr{)}\Bigr{)}=\exp\Bigl{(}\frac{n(\theta_{1}-% \theta_{0})(\theta_{0}+\theta_{1}-2\bar{x})}{2\sigma^{2}}\Bigr{)}.

Thus we have n=16n=16, σ=2\sigma=2 and x¯=10.3\bar{x}=10.3 and the exponent is 16⋅1⋅0.4/8=0.816\cdot 1\cdot 0.4/8=0.8 and the Bayes factor is e0.8=2.23e^{0.8}=2.23: the data are closer to the value 1010 than to 1111 and the data favour H0H_{0} by a factor of approximately two. If the prior odds are 11, the posterior odds are 2.232.23 and ℙ⁢(H0|x)=2.23/3.23=0.69\mathbb{P}(H_{0}\mskip 1.0mu|\mskip 1.0mux)=2.23/3.23=0.69. We see that the Bayes factor multiplies all prior opinions by the same factor, and in cases where the prior is very strong, it can be difficult to overturn the prior opinion. The situation in exercise 28.4 is more complicated, because the prior odds are not equal.

28.3 Comparing Frequentist and Bayesian Answers

Now we have two different sets of answers, for the same data, for the normal mean. We conclude by comparing the two sets.

The first difference is in the nature of the statements. A 95%95\% confidence interval is a statement about a procedure: it says that for repeated samples, the interval covers the fixed truth in 95%95\% of cases. A 95%95\% credible interval is a statement about the parameter, given the data and the prior. Similarly, a pp-value is computed assuming H0H_{0} is true, whereas ℙ⁢(H0|x)\mathbb{P}(H_{0}\mskip 1.0mu|\mskip 1.0mux) is the probability of the hypothesis given the data.

The second observation is that for a location parameter with a flat prior, both numbers coincide. Example 28.3 shows the two intervals coinciding, and the tests coincide as well: for H0:θ≤θ0H_{0}\colon\theta\leq\theta_{0} the flat-prior posterior probability Φ⁢(n⁢(θ0−x¯)/σ)\Phi\bigl{(}\sqrt{n}\,(\theta_{0}-\bar{x})/\sigma\bigr{)} equals, by the symmetry of Φ\Phi, the pp-value 1−Φ⁢(n⁢(x¯−θ0)/σ)1-\Phi\bigl{(}\sqrt{n}\,(\bar{x}-\theta_{0})/\sigma\bigr{)} of the one-sided zz-test from example 20.6; for the running data set both values equal 0.2740.274. This is the case because the flat-prior posterior density in θ\theta has the same shape as the sampling density of X¯\bar{X} in x¯\bar{x}. This is the reason why the misreadings in lectures 20 and 25 were so tempting: the forbidden sentence about the pp-value is, for the simplest model, a true sentence about the flat-prior posterior, and only the interpretation needs to be changed.

The coincidence is only true for special models and priors. A genuine prior will move the Bayesian answer closer to the prior mean and will concentrate the answer around this value, as in the case of N⁢(12,1)N(12,1) above. If the prior is correct, this is an advantage, but if the prior is wrong this effect will work against us. The frequentist guarantee from definition 25.1 applies for all θ\theta and thus cannot be violated in this way; exercise 28.5, however, shows that the frequentist coverage of a credible interval can be worse than 95%95\% when θ\theta is far from the prior mean.

As the sample size increases, the influence of the prior decreases. For the normal mean, equation (28.1) shows that the weight of the prior mean in mnm_{n} goes to zero and that τn2≤σ2/n\tau_{n}^{2}\leq\sigma^{2}/n so that both the posterior mean and the posterior standard deviation converge to the values they would have with flat prior. The same effect holds for all regular models. The general version of this result is the Bayesian analogue of theorem 17.1: under the regularity conditions from definition 16.1, and for any continuous, positive prior density at the true value, the posterior distribution is approximately

equation (28.4) (28.4)
N⁢(θ^n,1ℐ︀n⁢(θ^n))for large ⁢n,N\Bigl{(}\hat{\theta}_{n},\ \frac{1}{\mathcal{I}_{n}(\hat{\theta}_{n})}\Bigr{)% }\qquad\text{for large }n,

where θ^n\hat{\theta}_{n} is the maximum likelihood estimator. Without proof, we state that this result holds. The proof is similar to the one for theorem 17.1: the log-posterior is ℓ⁢(θ)+log⁡π⁢(θ)\ell(\theta)+\log\pi(\theta) up to a constant, the log-likelihood is approximately ℓ⁢(θ^n)−12⁢ℐ︀n⁢(θ^n)⁢(θ−θ^n)2\ell(\hat{\theta}_{n})-\tfrac{1}{2}\mathcal{I}_{n}(\hat{\theta}_{n})(\theta-% \hat{\theta}_{n})^{2} at its maximum, and this quadratic term grows like nn while log⁡π⁢(θ)\log\pi(\theta) does not. From (28.4) we know that the equal-tailed credible interval is approximately the Wald interval θ^n±z1−α/2⁢se(θ^n)\hat{\theta}_{n}\pm z_{1-\alpha/2}\,\mathop{\mathrm{se}}\nolimits(\hat{\theta}% _{n}) of proposition 25.7, independently of the prior: for large sample size both approaches will determine the same interval with approximately frequentist coverage 1−α1-\alpha.

The two approaches differ most when testing for a point null hypothesis H0:θ=θ0H_{0}\colon\theta=\theta_{0} against a two-sided alternative. In this case a Bayesian test requires a prior which puts a lump of probability on θ0\theta_{0} but which spreads the remaining probability over the alternative. The exact distribution of the remaining probability over the alternative matters: if the prior is diffuse on the alternative, H0H_{0} can have posterior probability of more than one half for a data set with pp-value 0.010.01. We will not consider this case in detail, but we can conclude that small pp-value and small posterior probability of H0H_{0} are different quantities and that in this case the difference can be substantial.

Summary.
  • •
    ​

    A credible interval with credibility 1−α1-\alpha has posterior probability 1−α1-\alpha: this is a probability statement about θ\theta for the given data, taking the prior into account. The equal-tailed interval is based on posterior quantiles, the HPD interval is the shortest one, and for symmetric, unimodal posteriors both intervals coincide.

  • •
    ​

    For the normal mean with known variance and normal prior, the 1−α1-\alpha credible interval is mn±z1−α/2⁢τnm_{n}\pm z_{1-\alpha/2}\,\tau_{n}; for the flat prior this interval coincides with the confidence interval x¯±z1−α/2⁢σ/n\bar{x}\pm z_{1-\alpha/2}\,\sigma/\sqrt{n}.

  • •
    ​

    A Bayesian test compares posterior probabilities of the hypotheses, or rejects H0H_{0} if ℙ⁢(H0|x)\mathbb{P}(H_{0}\mskip 1.0mu|\mskip 1.0mux) falls below a threshold. For two simple hypotheses, the posterior odds are the prior odds times the likelihood ratio, known as the Bayes factor.

  • •
    ​

    The level of a test and the confidence level of an interval are statements about the procedure applied to repeated samples, whereas the credibility of an interval is a statement about the parameter given one data set. For a location parameter with flat prior, the intervals, and the one-sided test and pp-value, agree numerically while the statements do not have the same meaning.

  • •
    ​

    An informative prior will shift and shorten the Bayesian answer; as n→∞n\to\infty the posterior is approximately N⁢(θ^n,1/ℐ︀n⁢(θ^n))N\bigl{(}\hat{\theta}_{n},1/\mathcal{I}_{n}(\hat{\theta}_{n})\bigr{)} and then the credible and Wald intervals coincide. For point-null hypothesis testing, the two approaches can differ significantly.

Exercise 28.1.

Let the posterior density π⁢(θ|x)\pi(\theta\mskip 1.0mu|\mskip 1.0mux) be continuous, strictly increasing on (−∞,μ](-\infty,\mu] and strictly decreasing on [μ,∞)[\mu,\infty).

  1. 1.
    ​

    Show that the HPD set HcH_{c} from definition 28.2 is an interval [l,u][l,u] with l<μ<ul<\mu<u and π⁢(l|x)=π⁢(u|x)\pi(l\mskip 1.0mu|\mskip 1.0mux)=\pi(u\mskip 1.0mu|\mskip 1.0mux).

  2. 2.
    ​

    Assume that the density is symmetric about μ\mu, i.e. that π⁢(μ+t|x)=π⁢(μ−t|x)\pi(\mu+t\mskip 1.0mu|\mskip 1.0mux)=\pi(\mu-t\mskip 1.0mu|\mskip 1.0mux) for all tt. Show that the HPD interval with credibility 1−α1-\alpha coincides with the equal-tailed interval [qα/2,q1−α/2][q_{\alpha/2},q_{1-\alpha/2}].

  3. 3.
    ​

    Now assume that the posterior density is continuous and strictly decreasing on [0,∞)[0,\infty) and zero elsewhere. Show that the HPD set with credibility 1−α1-\alpha is the one-sided interval [0,q1−α][0,q_{1-\alpha}] and that this interval is shorter than the equal-tailed interval.

Exercise 28.2.

In example 26.5, where a Bernoulli sample with n=10n=10 and T=7T=7 successes was combined with the prior Beta⁢(2,2)\text{Beta}(2,2) to obtain the posterior distribution Beta⁢(9,5)\text{Beta}(9,5) with mean 9/14=0.6439/14=0.643, the quantiles of the Beta distribution had to be found numerically. Using any statistical package, and using the qbeta function in R, we find the 0.0250.025- and 0.9750.975-quantiles of Beta⁢(9,5)\text{Beta}(9,5) to be 0.3860.386 and 0.8610.861.

  1. 1.
    ​

    Write down the equal-tailed 95%95\% credible interval for θ\theta and compare the mid-point of this interval to the posterior mean.

  2. 2.
    ​

    Show that the posterior density has its maximum at θ=2/3\theta=2/3, i.e. that the posterior distribution is skewed with the mode above the mean.

  3. 3.
    ​

    The HPD interval for Beta⁢(9,5)\text{Beta}(9,5), found numerically, is [0.401,0.874][0.401,0.874]. Using definition 28.2, explain why this interval is shifted towards the mode, compared to the equal-tailed interval. Also explain why the difference between the two intervals decreases as nn increases.

Exercise 28.3.

Assume that we have observed two successes in three Bernoulli trials with success probability θ\theta, and the prior distribution of θ\theta is flat on the interval (0,1)(0,1).

  1. 1.
    ​

    Determine the posterior distribution of θ\theta.

  2. 2.
    ​

    Calculate the posterior probability of H0:θ≤1/2H_{0}\colon\theta\leq 1/2 by integrating the posterior density. Also, calculate the posterior odds of H0H_{0} against H1:θ>1/2H_{1}\colon\theta>1/2.

  3. 3.
    ​

    The prior odds of H0H_{0} against H1H_{1} were 11. By what factor have the data changed these odds? How is this factor related to the likelihood? Why is it not simply the ratio of the likelihoods at the two parameter values, as in (28.3)?

  4. 4.
    ​

    Write down the equations for the boundaries of the 95%95\% equal-tailed credible interval and explain why these equations can only be solved numerically.

Exercise 28.4.

In the situation of example 20.6, a sample of size n=25n=25 from N⁢(θ,4)N(\theta,4) has mean x¯=10.9\bar{x}=10.9. We want to compare the simple hypotheses H0:θ=10H_{0}\colon\theta=10 and H1:θ=11H_{1}\colon\theta=11.

  1. 1.
    ​

    Compute the Bayes factor in favour of H0H_{0}.

  2. 2.
    ​

    Before the data were available, H0H_{0} was three times as likely as H1H_{1}. Compute the posterior odds and the posterior probability of H0H_{0}.

  3. 3.
    ​

    In lecture 20 we have seen that the one-sided pp-value 0.01220.0122 can be found for the data. Compare your two numbers and explain why they can disagree, despite the fact that both values are against H0H_{0}.

  4. 4.
    ​

    What value must the prior odds in favour of H0H_{0} have for the posterior probability of H0H_{0} to be greater than one half?

Exercise 28.5.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with N⁢(θ,σ2)N(\theta,\sigma^{2}) where σ2\sigma^{2} is known, and let the prior be N⁢(m,τ2)N(m,\tau^{2}). Let w=(1/τ2)/(1/τ2+n/σ2)w=(1/\tau^{2})/(1/\tau^{2}+n/\sigma^{2}) be the weight of the prior mean in mnm_{n}. Consider the credible interval (28.2) to be a random interval, as introduced in lecture 25, where θ\theta is fixed and the data are random.

  1. 1.
    ​

    Show that τn2=(1−w)⁢σ2/n\tau_{n}^{2}=(1-w)\sigma^{2}/n and that under ℙθ\mathbb{P}_{\theta} the posterior mean satisfies mn−θ∼N⁢(w⁢(m−θ),(1−w)2⁢σ2/n)m_{n}-\theta\sim N\bigl{(}w(m-\theta),\,(1-w)^{2}\sigma^{2}/n\bigr{)}.

  2. 2.
    ​

    Determine the coverage probability ℙθ⁢(mn−z1−α/2⁢τn≤θ≤mn+z1−α/2⁢τn)\mathbb{P}_{\theta}\bigl{(}m_{n}-z_{1-\alpha/2}\tau_{n}\leq\theta\leq m_{n}+z_% {1-\alpha/2}\tau_{n}\bigr{)} as a function of θ\theta.

  3. 3.
    ​

    For n=16n=16, σ=2\sigma=2, τ=1\tau=1, m=12m=12 and α=0.05\alpha=0.05, determine the coverage probability for θ=12\theta=12 and for θ=10\theta=10 and discuss the result as |θ−m|→∞|\theta-m|\to\infty.

  4. 4.
    ​

    Is the credible interval a confidence interval with level 0.950.95 in the sense of definition 25.1? What happens to the coverage probability as n→∞n\to\infty with θ\theta fixed?