Lecture 26 Bayesian Inference: Priors and Posteriors

So far the parameter θ\theta has been a fixed but unknown number, and a sentence such as “θ\theta is probably close to 0.30.3” had no meaning. In this lecture and in lecture 28 we will see that the parameter itself can be given a probability distribution: the prior, chosen before the data are seen, and the posterior, obtained after the data are seen. (We will see an update in pictures in lecture 27.)

26.1 Prior, Likelihood and Posterior

Since θ\theta is now a random variable, the model density is a conditional density given the parameter, and in the Bayesian lectures we write it as f⁢(x|θ)f(x\mskip 1.0mu|\mskip 1.0mu\theta), where x=(x1,…,xn)x=(x_{1},\dots,x_{n}) is the whole data set, so that f⁢(x|θ)=∏if⁢(xi|θ)f(x\mskip 1.0mu|\mskip 1.0mu\theta)=\prod_{i}f(x_{i}\mskip 1.0mu|\mskip 1.0mu\theta) is the likelihood L⁢(θ)L(\theta) of lecture 4.

Definition 26.1.

Let θ\theta be a random variable with values in Θ\Theta and density π⁢(θ)\pi(\theta), and assume that given θ\theta the data X=(X1,…,Xn)X=(X_{1},\dots,X_{n}) have joint density f⁢(x|θ)f(x\mskip 1.0mu|\mskip 1.0mu\theta). Then the density π\pi is called the prior density (or prior) of θ\theta and the conditional density π⁢(θ|x)\pi(\theta\mskip 1.0mu|\mskip 1.0mux) of θ\theta given X=xX=x is called the posterior density (or posterior). The parameters in the prior are called hyperparameters.

From now on, unless otherwise specified, π\pi will denote a prior or posterior density. In lecture 19, we used the same letter to denote the mixture weight for a two-component model.

The posterior density can be found using Bayes’ theorem as given in proposition A.20 in appendix A, where the parameter plays the role of AA and the data play the role of BB.

Theorem 26.2.

In the situation of definition 26.1, the posterior density is given by

equation (26.1) (26.1)
π⁢(θ|x)=f⁢(x|θ)⁢π⁢(θ)m⁢(x),where ⁢m⁢(x)=∫Θf⁢(x|ϑ)⁢π⁢(ϑ)⁢dϑ\pi(\theta\mskip 1.0mu|\mskip 1.0mux)=\frac{f(x\mskip 1.0mu|\mskip 1.0mu\theta% )\,\pi(\theta)}{m(x)},\qquad\text{where }m(x)=\int_{\Theta}f(x\mskip 1.0mu|% \mskip 1.0mu\vartheta)\,\pi(\vartheta)\,\mathrm{d}\vartheta

is the marginal density of the data. The formula holds for all xx with m⁢(x)>0m(x)>0.

Proof.

The joint density of (θ,X)(\theta,X) is given by f⁢(x|θ)⁢π⁢(θ)f(x\mskip 1.0mu|\mskip 1.0mu\theta)\,\pi(\theta). Integrating out θ\theta gives the marginal density m⁢(x)m(x) of XX. Dividing the joint density by m⁢(x)m(x) gives the conditional density of θ\theta given X=xX=x, as claimed. ∎

The denominator m⁢(x)m(x) makes the posterior integrate to one and does not involve θ\theta; in practice we never compute it directly, but write

equation (26.2) (26.2)
π⁢(θ|x)∝f⁢(x|θ)⁢π⁢(θ),\pi(\theta\mskip 1.0mu|\mskip 1.0mux)\propto f(x\mskip 1.0mu|\mskip 1.0mu% \theta)\,\pi(\theta),

where ∝\propto means “equal up to a factor free of θ\theta”, and we try to recognise the right-hand side as the kernel of a known density, whose normalising constant then comes from table A.2. In words, the posterior is proportional to the likelihood times the prior.

The posterior is the complete Bayesian answer: a statement like “ℙ⁢(θ>0.3|x)=0.95\mathbb{P}(\theta>0.3\mskip 1.0mu|\mskip 1.0mux)=0.95” now has a meaning. Point estimates are summaries of the posterior (section 26.3), and intervals and tests will be considered in lecture 28.

26.2 Conjugate Priors

The rule (26.2) works best if the posterior distribution belongs to the same family of distributions as the prior distribution, so that the update amounts to no more than a change of the hyperparameters.

Definition 26.3.

A family 𝒫︀\mathcal{P} of prior distributions is conjugate for a model (f⁢(x|θ)|θ∈Θ)\bigl{(}f(x\mskip 1.0mu|\mskip 1.0mu\theta)\!\mathrel{\big{|}}\!\theta\in% \Theta\bigr{)}, if for every prior in 𝒫︀\mathcal{P} and every data set xx the posterior π⁢(θ|x)\pi(\theta\mskip 1.0mu|\mskip 1.0mux) also belongs to 𝒫︀\mathcal{P}.

For exponential families, conjugate priors can be written down by inspection: the likelihood depends on the natural parameter η\eta only through the exponent η⁢Tn−n⁢K⁢(η)\eta T_{n}-nK(\eta) (lemma 7.5), and a prior with an exponent of the same form reproduces itself.

Proposition 26.4.

Let f⁢(x|η)=h⁢(x)⁢exp⁡(η⁢T⁢(x)−K⁢(η))f(x\mskip 1.0mu|\mskip 1.0mu\eta)=h(x)\exp\bigl{(}\eta T(x)-K(\eta)\bigr{)} with η∈𝒩︀\eta\in\mathcal{N} be a one-parameter exponential family in its natural parametrisation, and for real numbers aa and bb let

πa,b⁢(η)∝exp⁡(a⁢η−b⁢K⁢(η)),η∈𝒩︀,\pi_{a,b}(\eta)\propto\exp\bigl{(}a\eta-bK(\eta)\bigr{)},\qquad\eta\in\mathcal% {N},

whenever the right-hand side has a finite integral over 𝒩︀\mathcal{N}. Then these priors are conjugate for the model: if the prior is πa,b\pi_{a,b}, the data x1,…,xnx_{1},\dots,x_{n} have natural statistic Tn=∑i=1nT⁢(xi)T_{n}=\sum_{i=1}^{n}T(x_{i}) and the marginal density satisfies 0<m⁢(x)<∞0<m(x)<\infty, then the posterior is πa+Tn,b+n\pi_{a+T_{n},\,b+n}.

Proof.

Using (26.2) and lemma 7.5 we find

π⁢(η|x)\displaystyle\pi(\eta\mskip 1.0mu|\mskip 1.0mux)
∝∏i=1nh⁢(xi)⁢exp⁡(η⁢Tn−n⁢K⁢(η))⁢exp⁡(a⁢η−b⁢K⁢(η))\displaystyle\propto\prod_{i=1}^{n}h(x_{i})\,\exp\bigl{(}\eta T_{n}-nK(\eta)% \bigr{)}\,\exp\bigl{(}a\eta-bK(\eta)\bigr{)}
∝exp⁡((a+Tn)⁢η−(b+n)⁢K⁢(η)),\displaystyle\propto\exp\bigl{(}(a+T_{n})\eta-(b+n)K(\eta)\bigr{)},

since ∏ih⁢(xi)\prod_{i}h(x_{i}) does not depend on η\eta. The right-hand side is the kernel of πa+Tn,b+n\pi_{a+T_{n},b+n} and by assumption m⁢(x)<∞m(x)<\infty this kernel has finite integral and thus is a density. This shows that the posterior is πa+Tn,b+n\pi_{a+T_{n},b+n}. This completes the proof. ∎

The update adds the natural statistic to aa and the sample size to bb: the prior acts like bb earlier observations with statistic total aa. We now work the recipe for our simplest model in the original parameter.

Example 26.5.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Bernoulli with success probability θ∈(0,1)\theta\in(0,1) and let the prior be the Beta⁢(a,b)\text{Beta}(a,b) distribution with a,b>0a,b>0. From table A.2 we know that this distribution has density

π⁢(θ)=θa−1⁢(1−θ)b−1B⁢(a,b),0<θ<1,\pi(\theta)=\frac{\theta^{a-1}(1-\theta)^{b-1}}{B(a,b)},\qquad 0<\theta<1,

and the mean of this distribution is a/(a+b)a/(a+b). If T=∑ixiT=\sum_{i}x_{i} is the number of successes in the sample, then the likelihood is f⁢(x|θ)=θT⁢(1−θ)n−Tf(x\mskip 1.0mu|\mskip 1.0mu\theta)=\theta^{T}(1-\theta)^{n-T}, as in example 4.2, and thus we have

π⁢(θ|x)∝θT⁢(1−θ)n−T⋅θa−1⁢(1−θ)b−1=θa+T−1⁢(1−θ)b+n−T−1,\pi(\theta\mskip 1.0mu|\mskip 1.0mux)\propto\theta^{T}(1-\theta)^{n-T}\cdot% \theta^{a-1}(1-\theta)^{b-1}=\theta^{a+T-1}(1-\theta)^{b+n-T-1},

where we have omitted the constant 1/B⁢(a,b)1/B(a,b) for simplicity. The right-hand side is the kernel of a Beta⁢(a+T,b+n−T)\text{Beta}(a+T,\,b+n-T) density and thus we have found the posterior distribution: the Beta distributions are conjugate for the Bernoulli model, and the update adds the successes to aa and the failures to bb. The posterior mean is given by

equation (26.3) (26.3)
𝔼⁢(θ|x)=a+Ta+b+n=a+ba+b+n⋅aa+b+na+b+n⋅x¯,\mathbb{E}(\theta\mskip 1.0mu|\mskip 1.0mux)=\frac{a+T}{a+b+n}=\frac{a+b}{a+b+% n}\cdot\frac{a}{a+b}+\frac{n}{a+b+n}\cdot\bar{x},

which is a weighted average of the prior mean a/(a+b)a/(a+b) and the sample proportion x¯=T/n\bar{x}=T/n, i.e. the maximum likelihood estimate from example 5.3. The prior can be interpreted as having observed a+ba+b additional samples, of which aa were successes. This equivalent sample size is a measure for how the hyperparameters should be chosen in practice (exercise 26.1). As n→∞n\to\infty, the weight of the prior decreases and with enough data the prior will be completely forgotten. For a=b=2a=b=2, n=10n=10 and T=7T=7, for example, the posterior distribution is Beta⁢(9,5)\text{Beta}(9,5) with mean 9/14≈0.6439/14\approx 0.643, between the prior mean 0.50.5 and the sample proportion 0.70.7.

In terms of the log-odds η\eta, the example is an instance of proposition 26.4. Using the same arguments, we can show that normal priors are conjugate for the normal mean with known variance. The cumulant function in this case is a quadratic function of η\eta (exercise 7.4). The result is given here without proof, and the proof, based on completing the square, is given as exercise 26.3.

Proposition 26.6.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. N⁢(θ,σ2)N(\theta,\sigma^{2}) with known variance σ2\sigma^{2} and let the prior be θ∼N⁢(m,τ2)\theta\sim N(m,\tau^{2}). Then the posterior distribution is N⁢(mn,τn2)N(m_{n},\tau_{n}^{2}) where

1τn2=1τ2+nσ2andmn=1/τ21/τ2+n/σ2⁢m+n/σ21/τ2+n/σ2⁢x¯.\frac{1}{\tau_{n}^{2}}=\frac{1}{\tau^{2}}+\frac{n}{\sigma^{2}}\qquad\text{and}% \qquad m_{n}=\frac{1/\tau^{2}}{1/\tau^{2}+n/\sigma^{2}}\,m+\frac{n/\sigma^{2}}% {1/\tau^{2}+n/\sigma^{2}}\,\bar{x}.

The reciprocal of a variance is called a precision, and in these terms the proposition is easy to remember: precisions add, and the posterior mean is the precision-weighted average of the prior mean and x¯\bar{x}. Again, the prior washes out as n→∞n\to\infty.

26.3 Point Estimates from the Posterior

If a single number is required for θ\theta, the posterior mean, median and mode can all be used. If the errors are measured in terms of squared error, as in lecture 2, the posterior mean is the optimal choice.

Theorem 26.7.

Assume that the posterior distribution has finite variance. Then the posterior mean θ^B=𝔼⁢(θ|x)\hat{\theta}_{B}=\mathbb{E}(\theta\mskip 1.0mu|\mskip 1.0mux) minimises the posterior expected squared error, i.e. for every number dd we have

𝔼⁢((θ−d)2|x)≥𝔼⁢((θ−θ^B)2|x)=Var(θ|x),\mathbb{E}\bigl{(}(\theta-d)^{2}\!\mathrel{\big{|}}\!x\bigr{)}\geq\mathbb{E}% \bigl{(}(\theta-\hat{\theta}_{B})^{2}\!\mathrel{\big{|}}\!x\bigr{)}=\mathop{% \mathrm{Var}}\nolimits(\theta\mskip 1.0mu|\mskip 1.0mux),

and the equality holds only for d=θ^Bd=\hat{\theta}_{B}.

Proof.

Adding and subtracting θ^B\hat{\theta}_{B} inside the square, as in the proof of theorem 2.7, we get

𝔼⁢((θ−d)2|x)=𝔼⁢((θ−θ^B)2|x)+2⁢(θ^B−d)⁢𝔼⁢(θ−θ^B|x)+(θ^B−d)2,\mathbb{E}\bigl{(}(\theta-d)^{2}\!\mathrel{\big{|}}\!x\bigr{)}=\mathbb{E}\bigl% {(}(\theta-\hat{\theta}_{B})^{2}\!\mathrel{\big{|}}\!x\bigr{)}+2(\hat{\theta}_% {B}-d)\,\mathbb{E}(\theta-\hat{\theta}_{B}\mskip 1.0mu|\mskip 1.0mux)+(\hat{% \theta}_{B}-d)^{2},

where all expectations are taken with respect to the posterior. The middle term vanishes, because 𝔼⁢(θ|x)=θ^B\mathbb{E}(\theta\mskip 1.0mu|\mskip 1.0mux)=\hat{\theta}_{B}, and the last term is non-negative and zero only for d=θ^Bd=\hat{\theta}_{B}. This completes the proof. ∎

The estimator θ^B\hat{\theta}_{B} is called the Bayes estimator (under squared error) for θ\theta. The theorem is a statement about fixed data. For fixed θ\theta, and varying data, exercise 26.4 shows that the Bernoulli θ^B\hat{\theta}_{B} is biased but consistent, and that neither the Bernoulli estimator nor X¯\bar{X} has smaller mean squared error for all θ\theta.

26.4 Flat, Improper and Jeffreys Priors

Without prior knowledge to encode, the obvious choice is the flat prior which is a constant density on Θ\Theta. This prior does not prefer any value of θ\theta. For the Bernoulli model this prior is Beta⁢(1,1)\text{Beta}(1,1) and from example 26.5 we know that the posterior mean is (T+1)/(n+2)(T+1)/(n+2), which is close to x¯\bar{x} but is shifted towards 1/21/2. Since a constant density on an unbounded parameter space has infinite integral, the flat prior is improper. Nevertheless, this prior is still in use today: if f⁢(x|θ)⁢π⁢(θ)f(x\mskip 1.0mu|\mskip 1.0mu\theta)\,\pi(\theta) has finite integral, it is normalised and used as the posterior. The finiteness of the integral must be checked in each case (exercise 26.2). For the normal mean, the flat prior on ℝ\mathbb{R} is the limit τ→∞\tau\to\infty in proposition 26.6 and in this case leads to the proper posterior N⁢(x¯,σ2/n)N(\bar{x},\sigma^{2}/n).

The flat prior has a more basic defect: flatness depends on the parametrisation. If θ\theta is uniform on (0,1)(0,1), then ψ=θ2\psi=\theta^{2} has density 1/(2⁢ψ)1/(2\sqrt{\psi}) by proposition A.19, and a prior which claims to know nothing about θ\theta claims to know something about θ2\theta^{2}. The Fisher information provides a parametrisation-free rule, since proposition 11.10 shows how it transforms.

Definition 26.8.

In the situation of section 11.1, the Jeffreys prior for θ\theta is given by

πJ⁢(θ)∝ℐ︀⁢(θ),θ∈Θ,\pi_{J}(\theta)\propto\sqrt{\mathcal{I}(\theta)},\qquad\theta\in\Theta,

where ℐ︀⁢(θ)\mathcal{I}(\theta) is the Fisher information of an observation.

This prior is justified by the following invariance property, which the flat prior lacks.

Proposition 26.9.

Let ψ=g⁢(θ)\psi=g(\theta) where g:Θ→g⁢(Θ)g\colon\Theta\to g(\Theta) is a continuously differentiable bijection with g′⁢(θ)≠0g^{\prime}(\theta)\neq 0 for all θ\theta and let h=g−1h=g^{-1}.

  1. 1.
    ​

    If θ\theta has the flat prior on Θ\Theta, then ψ\psi has the density |h′⁢(ψ)||h^{\prime}(\psi)| on g⁢(Θ)g(\Theta). The density is flat only if gg is affine.

  2. 2.
    ​

    If θ\theta has the Jeffreys prior ℐ︀θ⁢(θ)\sqrt{\mathcal{I}_{\theta}(\theta)}, then ψ\psi has the density ℐ︀ψ⁢(ψ)\sqrt{\mathcal{I}_{\psi}(\psi)}. This is the Jeffreys prior computed using the ψ\psi-parametrisation.

Proof.

Using the change-of-variable formula from proposition A.19, if a density πθ\pi_{\theta} for θ\theta is given, then a density of πψ⁢(ψ)=πθ⁢(h⁢(ψ))⁢|h′⁢(ψ)|\pi_{\psi}(\psi)=\pi_{\theta}\bigl{(}h(\psi)\bigr{)}\,|h^{\prime}(\psi)| for ψ\psi can be derived. The same substitution can be applied to improper densities. In the case of the flat prior πθ≡1\pi_{\theta}\equiv 1, we have |h′⁢(ψ)||h^{\prime}(\psi)| and the derivative of hh is constant only if hh, i.e. if gg, is affine. For the Jeffreys prior we have πθ⁢(θ)=ℐ︀θ⁢(θ)\pi_{\theta}(\theta)=\sqrt{\mathcal{I}_{\theta}(\theta)} and using proposition 11.10 and h′⁢(ψ)=1/g′⁢(θ)h^{\prime}(\psi)=1/g^{\prime}(\theta) for θ=h⁢(ψ)\theta=h(\psi) we get

πψ⁢(ψ)=ℐ︀θ⁢(θ)⁢|h′⁢(ψ)|=ℐ︀θ⁢(θ)|g′⁢(θ)|=ℐ︀θ⁢(θ)g′⁢(θ)2=ℐ︀ψ⁢(ψ).\pi_{\psi}(\psi)=\sqrt{\mathcal{I}_{\theta}(\theta)}\,\bigl{|}h^{\prime}(\psi)% \bigr{|}=\frac{\sqrt{\mathcal{I}_{\theta}(\theta)}}{|g^{\prime}(\theta)|}=% \sqrt{\frac{\mathcal{I}_{\theta}(\theta)}{g^{\prime}(\theta)^{2}}}=\sqrt{% \mathcal{I}_{\psi}(\psi)}.

This is the Jeffreys prior in the ψ\psi-parametrisation. This completes the proof. ∎

The Jeffreys prior is a property of the model, rather than of the way it is labelled, and as such it can be proper or improper.

Example 26.10.

For the Bernoulli model, from example 11.5 we know ℐ︀⁢(θ)=1/(θ⁢(1−θ))\mathcal{I}(\theta)=1/\bigl{(}\theta(1-\theta)\bigr{)} and thus the Jeffreys prior is given by

πJ⁢(θ)∝θ−1/2⁢(1−θ)−1/2,0<θ<1.\pi_{J}(\theta)\propto\theta^{-1/2}(1-\theta)^{-1/2},\qquad 0<\theta<1.

This is the density of the Beta⁢(1/2,1/2)\text{Beta}(1/2,1/2) distribution. The Jeffreys prior is proper and has more weight near the boundaries of the interval than the flat prior does. From example 26.5 we know that the posterior distribution is Beta⁢(T+1/2,n−T+1/2)\text{Beta}(T+1/2,\,n-T+1/2) and the mean of the posterior is (T+1/2)/(n+1)(T+1/2)/(n+1). This value is between the maximum likelihood estimate T/nT/n and the estimate (T+1)/(n+2)(T+1)/(n+2) obtained using the flat prior. For large nn, all three values coincide.

For location and scale models the Jeffreys prior is the same for every base density f0f_{0}.

Proposition 26.11.

Let f0f_{0} be a fixed density such that the following models satisfy the conditions of section 11.1.

  1. 1.
    ​

    Consider the location model f⁢(x|θ)=f0⁢(x−θ)f(x\mskip 1.0mu|\mskip 1.0mu\theta)=f_{0}(x-\theta) with θ∈ℝ\theta\in\mathbb{R}. The Fisher information in this model is constant as a function of θ\theta and thus the Jeffreys prior is flat, πJ⁢(θ)∝1\pi_{J}(\theta)\propto 1 on ℝ\mathbb{R}.

  2. 2.
    ​

    Consider the scale model f⁢(x|σ)=f0⁢(x/σ)/σf(x\mskip 1.0mu|\mskip 1.0mu\sigma)=f_{0}(x/\sigma)/\sigma with σ>0\sigma>0. The Fisher information in this model is ℐ︀⁢(σ)=c/σ2\mathcal{I}(\sigma)=c/\sigma^{2} for a constant c>0c>0. Thus, the Jeffreys prior is πJ⁢(σ)∝1/σ\pi_{J}(\sigma)\propto 1/\sigma on (0,∞)(0,\infty).

Both priors are improper.

Exercise 26.5 verifies both statements by writing the score as a function of the standardised observation. The normal mean with known variance is a location model, so its Jeffreys prior is flat. Similarly, the normal standard deviation with known mean is a scale model with ℐ︀⁢(σ)=2/σ2\mathcal{I}(\sigma)=2/\sigma^{2}.

Summary.
  • •
    ​

    The data can be used to update the prior density π⁢(θ)\pi(\theta) to the posterior π⁢(θ|x)∝f⁢(x|θ)⁢π⁢(θ)\pi(\theta\mskip 1.0mu|\mskip 1.0mux)\propto f(x\mskip 1.0mu|\mskip 1.0mu% \theta)\,\pi(\theta). The normalising constant can be found by recognising the kernel of a known density.

  • •
    ​

    Priors are conjugate for a family if the posterior always belongs to the same family. For a one-parameter exponential family in its natural parametrisation, the priors exp⁡(a⁢η−b⁢K⁢(η))\exp(a\eta-bK(\eta)) are conjugate. Beta priors are conjugate for the Bernoulli model, and normal priors are conjugate for the normal mean with known variance.

  • •
    ​

    The posterior mean is a weighted average of the prior mean and the maximum likelihood estimate. As nn increases, the prior weight decreases. The posterior mean is the Bayes estimator for θ\theta under squared error.

  • •
    ​

    A flat prior is not invariant under reparametrisation, but the Jeffreys prior πJ⁢(θ)∝ℐ︀⁢(θ)\pi_{J}(\theta)\propto\sqrt{\mathcal{I}(\theta)} is. The Jeffreys prior is Beta⁢(1/2,1/2)\text{Beta}(1/2,1/2) for the Bernoulli model, is flat for a location parameter, and is proportional to 1/σ1/\sigma for a scale parameter. An improper prior can only be used if the posterior is proper.

Exercise 26.1.

A Bernoulli model is to be fitted with a Beta⁢(a,b)\text{Beta}(a,b) prior for the success probability θ\theta.

  1. 1.
    ​

    Find aa and bb such that the prior mean is 0.30.3 and the equivalent sample size a+ba+b is 1010. You can find the prior standard deviation in table A.2.

  2. 2.
    ​

    Find aa and bb such that the prior has mean 0.30.3 and standard deviation 0.10.1. What is the equivalent sample size for this prior?

  3. 3.
    ​

    For the prior in the first part of the question, a sample of size n=20n=20 has T=9T=9 successes. Determine the posterior distribution and the posterior mean. Write the posterior mean as a weighted average of the prior mean and x¯\bar{x}.

Exercise 26.2.

The Haldane prior for the Bernoulli success probability is the limit a,b→0a,b\to 0 of the Beta⁢(a,b)\text{Beta}(a,b) prior, i.e. the function π⁢(θ)∝θ−1⁢(1−θ)−1\pi(\theta)\propto\theta^{-1}(1-\theta)^{-1} for 0<θ<10<\theta<1.

  1. 1.
    ​

    Show that this prior is improper.

  2. 2.
    ​

    Show that the formal posterior π⁢(θ|x)∝θT−1⁢(1−θ)n−T−1\pi(\theta\mskip 1.0mu|\mskip 1.0mux)\propto\theta^{T-1}(1-\theta)^{n-T-1} is a proper distribution if and only if 0<T<n0<T<n. Identify the resulting distribution.

  3. 3.
    ​

    For 0<T<n0<T<n, determine the posterior mean and compare it to the maximum likelihood estimate.

  4. 4.
    ​

    Describe what happens for T=0T=0.

Exercise 26.3.

Using the following steps, prove proposition 26.6.

  1. 1.
    ​

    Show that the likelihood, as a function of θ\theta, satisfies f⁢(x|θ)∝exp⁡(−n⁢(θ−x¯)2/(2⁢σ2))f(x\mskip 1.0mu|\mskip 1.0mu\theta)\propto\exp\bigl{(}-n(\theta-\bar{x})^{2}/(% 2\sigma^{2})\bigr{)}.

  2. 2.
    ​

    Multiply the prior density with this function and show, by completing the square in θ\theta, that

    π⁢(θ|x)∝exp⁡(−(θ−mn)22⁢τn2)\pi(\theta\mskip 1.0mu|\mskip 1.0mux)\propto\exp\Bigl{(}-\frac{(\theta-m_{n})^% {2}}{2\tau_{n}^{2}}\Bigr{)}

    where mnm_{n} and τn2\tau_{n}^{2} are as in the proposition.

  3. 3.
    ​

    What happens to mnm_{n} and τn2\tau_{n}^{2} when τ→∞\tau\to\infty with nn fixed? And when n→∞n\to\infty with τ\tau fixed?

Exercise 26.4.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. Bernoulli with success probability θ\theta and let θ^B=(a+T)/(a+b+n)\hat{\theta}_{B}=(a+T)/(a+b+n) with T=∑iXiT=\sum_{i}X_{i} be the Bayes estimator from example 26.5. Now fix θ\theta and let θ^B\hat{\theta}_{B} be an estimator as in lecture 2.

  1. 1.
    ​

    What are the bias and variance of θ^B\hat{\theta}_{B}? Show that the estimator is biased for every θ≠a/(a+b)\theta\neq a/(a+b), but consistent.

  2. 2.
    ​

    For a=b=1a=b=1 and n=10n=10, compare the estimators MSE(θ^B)\mathop{\mathrm{MSE}}\nolimits(\hat{\theta}_{B}) and MSE(X¯)\mathop{\mathrm{MSE}}\nolimits(\bar{X}) for θ=1/2\theta=1/2 and θ=0.1\theta=0.1.

Exercise 26.5.

Let f0f_{0} be a fixed density as in proposition 26.11 and let ℓ0⁢(z)=log⁡f0⁢(z)\ell_{0}(z)=\log f_{0}(z).

  1. 1.
    ​

    For the location model f⁢(x|θ)=f0⁢(x−θ)f(x\mskip 1.0mu|\mskip 1.0mu\theta)=f_{0}(x-\theta), show that the score is −ℓ0′⁢(x−θ)-\ell_{0}^{\prime}(x-\theta) and thus that the Fisher information does not depend on θ\theta.

  2. 2.
    ​

    For the scale model f⁢(x|σ)=f0⁢(x/σ)/σf(x\mskip 1.0mu|\mskip 1.0mu\sigma)=f_{0}(x/\sigma)/\sigma, show that the score is −(1+z⁢ℓ0′⁢(z))/σ-\bigl{(}1+z\ell_{0}^{\prime}(z)\bigr{)}/\sigma where z=x/σz=x/\sigma and that ℐ︀⁢(σ)=c/σ2\mathcal{I}(\sigma)=c/\sigma^{2} where cc does not depend on σ\sigma.

  3. 3.
    ​

    For the N⁢(0,σ2)N(0,\sigma^{2}) model with parameter σ\sigma, show that ℐ︀⁢(σ)=2/σ2\mathcal{I}(\sigma)=2/\sigma^{2}.