Lecture 26 Bayesian Inference: Priors and Posteriors
So far the parameter has been a fixed but unknown number, and a sentence such as “ is probably close to ” had no meaning. In this lecture and in lecture 28 we will see that the parameter itself can be given a probability distribution: the prior, chosen before the data are seen, and the posterior, obtained after the data are seen. (We will see an update in pictures in lecture 27.)
26.1 Prior, Likelihood and Posterior
Since is now a random variable, the model density is a conditional density given the parameter, and in the Bayesian lectures we write it as , where is the whole data set, so that is the likelihood of lecture 4.
Let be a random variable with values in and density , and assume that given the data have joint density . Then the density is called the prior density (or prior) of and the conditional density of given is called the posterior density (or posterior). The parameters in the prior are called hyperparameters.
From now on, unless otherwise specified, will denote a prior or posterior density. In lecture 19, we used the same letter to denote the mixture weight for a two-component model.
The posterior density can be found using Bayes’ theorem as given in proposition A.20 in appendix A, where the parameter plays the role of and the data play the role of .
In the situation of definition 26.1, the posterior density is given by
is the marginal density of the data. The formula holds for all with .
The joint density of is given by . Integrating out gives the marginal density of . Dividing the joint density by gives the conditional density of given , as claimed. ∎
The denominator makes the posterior integrate to one and does not involve ; in practice we never compute it directly, but write
where means “equal up to a factor free of ”, and we try to recognise the right-hand side as the kernel of a known density, whose normalising constant then comes from table A.2. In words, the posterior is proportional to the likelihood times the prior.
26.2 Conjugate Priors
The rule (26.2) works best if the posterior distribution belongs to the same family of distributions as the prior distribution, so that the update amounts to no more than a change of the hyperparameters.
A family of prior distributions is conjugate for a model , if for every prior in and every data set the posterior also belongs to .
For exponential families, conjugate priors can be written down by inspection: the likelihood depends on the natural parameter only through the exponent (lemma 7.5), and a prior with an exponent of the same form reproduces itself.
Let with be a one-parameter exponential family in its natural parametrisation, and for real numbers and let
whenever the right-hand side has a finite integral over . Then these priors are conjugate for the model: if the prior is , the data have natural statistic and the marginal density satisfies , then the posterior is .
The update adds the natural statistic to and the sample size to : the prior acts like earlier observations with statistic total . We now work the recipe for our simplest model in the original parameter.
Let be i.i.d. Bernoulli with success probability and let the prior be the distribution with . From table A.2 we know that this distribution has density
and the mean of this distribution is . If is the number of successes in the sample, then the likelihood is , as in example 4.2, and thus we have
where we have omitted the constant for simplicity. The right-hand side is the kernel of a density and thus we have found the posterior distribution: the Beta distributions are conjugate for the Bernoulli model, and the update adds the successes to and the failures to . The posterior mean is given by
which is a weighted average of the prior mean and the sample proportion , i.e. the maximum likelihood estimate from example 5.3. The prior can be interpreted as having observed additional samples, of which were successes. This equivalent sample size is a measure for how the hyperparameters should be chosen in practice (exercise 26.1). As , the weight of the prior decreases and with enough data the prior will be completely forgotten. For , and , for example, the posterior distribution is with mean , between the prior mean and the sample proportion .
In terms of the log-odds , the example is an instance of proposition 26.4. Using the same arguments, we can show that normal priors are conjugate for the normal mean with known variance. The cumulant function in this case is a quadratic function of (exercise 7.4). The result is given here without proof, and the proof, based on completing the square, is given as exercise 26.3.
Let be i.i.d. with known variance and let the prior be . Then the posterior distribution is where
The reciprocal of a variance is called a precision, and in these terms the proposition is easy to remember: precisions add, and the posterior mean is the precision-weighted average of the prior mean and . Again, the prior washes out as .
26.3 Point Estimates from the Posterior
If a single number is required for , the posterior mean, median and mode can all be used. If the errors are measured in terms of squared error, as in lecture 2, the posterior mean is the optimal choice.
Assume that the posterior distribution has finite variance. Then the posterior mean minimises the posterior expected squared error, i.e. for every number we have
and the equality holds only for .
Adding and subtracting inside the square, as in the proof of theorem 2.7, we get
where all expectations are taken with respect to the posterior. The middle term vanishes, because , and the last term is non-negative and zero only for . This completes the proof. ∎
The estimator is called the Bayes estimator (under squared error) for . The theorem is a statement about fixed data. For fixed , and varying data, exercise 26.4 shows that the Bernoulli is biased but consistent, and that neither the Bernoulli estimator nor has smaller mean squared error for all .
26.4 Flat, Improper and Jeffreys Priors
Without prior knowledge to encode, the obvious choice is the flat prior which is a constant density on . This prior does not prefer any value of . For the Bernoulli model this prior is and from example 26.5 we know that the posterior mean is , which is close to but is shifted towards . Since a constant density on an unbounded parameter space has infinite integral, the flat prior is improper. Nevertheless, this prior is still in use today: if has finite integral, it is normalised and used as the posterior. The finiteness of the integral must be checked in each case (exercise 26.2). For the normal mean, the flat prior on is the limit in proposition 26.6 and in this case leads to the proper posterior .
The flat prior has a more basic defect: flatness depends on the parametrisation. If is uniform on , then has density by proposition A.19, and a prior which claims to know nothing about claims to know something about . The Fisher information provides a parametrisation-free rule, since proposition 11.10 shows how it transforms.
In the situation of section 11.1, the Jeffreys prior for is given by
where is the Fisher information of an observation.
This prior is justified by the following invariance property, which the flat prior lacks.
Let where is a continuously differentiable bijection with for all and let .
-
1.
If has the flat prior on , then has the density on . The density is flat only if is affine.
-
2.
If has the Jeffreys prior , then has the density . This is the Jeffreys prior computed using the -parametrisation.
Using the change-of-variable formula from proposition A.19, if a density for is given, then a density of for can be derived. The same substitution can be applied to improper densities. In the case of the flat prior , we have and the derivative of is constant only if , i.e. if , is affine. For the Jeffreys prior we have and using proposition 11.10 and for we get
This is the Jeffreys prior in the -parametrisation. This completes the proof. ∎
The Jeffreys prior is a property of the model, rather than of the way it is labelled, and as such it can be proper or improper.
For the Bernoulli model, from example 11.5 we know and thus the Jeffreys prior is given by
This is the density of the distribution. The Jeffreys prior is proper and has more weight near the boundaries of the interval than the flat prior does. From example 26.5 we know that the posterior distribution is and the mean of the posterior is . This value is between the maximum likelihood estimate and the estimate obtained using the flat prior. For large , all three values coincide.
For location and scale models the Jeffreys prior is the same for every base density .
Let be a fixed density such that the following models satisfy the conditions of section 11.1.
-
1.
Consider the location model with . The Fisher information in this model is constant as a function of and thus the Jeffreys prior is flat, on .
-
2.
Consider the scale model with . The Fisher information in this model is for a constant . Thus, the Jeffreys prior is on .
Both priors are improper.
Exercise 26.5 verifies both statements by writing the score as a function of the standardised observation. The normal mean with known variance is a location model, so its Jeffreys prior is flat. Similarly, the normal standard deviation with known mean is a scale model with .
-
•
The data can be used to update the prior density to the posterior . The normalising constant can be found by recognising the kernel of a known density.
-
•
Priors are conjugate for a family if the posterior always belongs to the same family. For a one-parameter exponential family in its natural parametrisation, the priors are conjugate. Beta priors are conjugate for the Bernoulli model, and normal priors are conjugate for the normal mean with known variance.
-
•
The posterior mean is a weighted average of the prior mean and the maximum likelihood estimate. As increases, the prior weight decreases. The posterior mean is the Bayes estimator for under squared error.
-
•
A flat prior is not invariant under reparametrisation, but the Jeffreys prior is. The Jeffreys prior is for the Bernoulli model, is flat for a location parameter, and is proportional to for a scale parameter. An improper prior can only be used if the posterior is proper.
A Bernoulli model is to be fitted with a prior for the success probability .
-
1.
Find and such that the prior mean is and the equivalent sample size is . You can find the prior standard deviation in table A.2.
-
2.
Find and such that the prior has mean and standard deviation . What is the equivalent sample size for this prior?
-
3.
For the prior in the first part of the question, a sample of size has successes. Determine the posterior distribution and the posterior mean. Write the posterior mean as a weighted average of the prior mean and .
The Haldane prior for the Bernoulli success probability is the limit of the prior, i.e. the function for .
-
1.
Show that this prior is improper.
-
2.
Show that the formal posterior is a proper distribution if and only if . Identify the resulting distribution.
-
3.
For , determine the posterior mean and compare it to the maximum likelihood estimate.
-
4.
Describe what happens for .
Using the following steps, prove proposition 26.6.
-
1.
Show that the likelihood, as a function of , satisfies .
-
2.
Multiply the prior density with this function and show, by completing the square in , that
where and are as in the proposition.
-
3.
What happens to and when with fixed? And when with fixed?
Let be i.i.d. Bernoulli with success probability and let with be the Bayes estimator from example 26.5. Now fix and let be an estimator as in lecture 2.
-
1.
What are the bias and variance of ? Show that the estimator is biased for every , but consistent.
-
2.
For and , compare the estimators and for and .
Let be a fixed density as in proposition 26.11 and let .
-
1.
For the location model , show that the score is and thus that the Fisher information does not depend on .
-
2.
For the scale model , show that the score is where and that where does not depend on .
-
3.
For the model with parameter , show that .