Appendix C Solutions to the Exercises

This appendix contains worked solutions to the exercises found at the end of each lecture. The solutions are grouped by lecture and refer to the exercises by number.

C.1 Statistical Models and the Estimation Problem

Here we solve the exercises at the end of lecture 1.

Solution to exercise 1.1.
  1. 1.
    ​

    Since the tosses are independent and since each toss has probability θ\theta of being heads, we can describe X1,…,XnX_{1},\dots,X_{n} as a sequence of independent and identically distributed Bernoulli trials with probability weights f⁢(x;θ)=θx⁢(1−θ)1−xf(x;\theta)=\theta^{x}(1-\theta)^{1-x} for x∈{0,1}x\in\{0,1\} and parameter space Θ=(0,1)\Theta=(0,1).

  2. 2.
    ​

    We can use the proportion of heads in the sample, θ^=X¯=(1/n)⁢∑i=1nXi\hat{\theta}=\bar{X}=(1/n)\sum_{i=1}^{n}X_{i}, as an estimator.

  3. 3.
    ​

    The value θ^⁢(X1,…,Xn)\hat{\theta}(X_{1},\dots,X_{n}) is a function of the random variables XiX_{i} before the coin tosses occur, so the value is still indeterminate. After the tosses have been made and the numbers x1,…,xnx_{1},\dots,x_{n} are known, θ^⁢(x1,…,xn)\hat{\theta}(x_{1},\dots,x_{n}) is a function of numbers, i.e. a single number.

∎

Solution to exercise 1.2.

A statistic is a function of the data which does not use the unknown parameter θ=(μ,σ2)\theta=(\mu,\sigma^{2}).

  1. 1.
    ​

    The range maxi⁡Xi−mini⁡Xi\max_{i}X_{i}-\min_{i}X_{i} is computed from the data alone, so it is a statistic.

  2. 2.
    ​

    The value (1/n)⁢∑i(Xi−μ)2(1/n)\sum_{i}(X_{i}-\mu)^{2} contains the unknown mean μ\mu and thus cannot be computed from the data, so it is not a statistic. If μ\mu was known, then it would be a statistic.

  3. 3.
    ​

    The sample variance (1/(n−1))⁢∑i(Xi−X¯)2\bigl{(}1/(n-1)\bigr{)}\sum_{i}(X_{i}-\bar{X})^{2} uses the sample mean X¯\bar{X}, which is computed from the data, in place of μ\mu. Thus, it is a statistic.

∎

Solution to exercise 1.3.
  1. 1.
    ​

    Since the observations are independent, the joint density is the product of the marginal densities:

    f⁢(x1,…,xn;θ)=∏i=1n1θ=1θn,f(x_{1},\dots,x_{n};\theta)=\prod_{i=1}^{n}\frac{1}{\theta}=\frac{1}{\theta^{% \,n}},

    whenever all observations are in the support, i.e. when 0≤xi≤θ0\leq x_{i}\leq\theta for all ii. Otherwise, the density is zero. This can also be written as the density being non-zero only when mini⁡xi≥0\min_{i}x_{i}\geq 0 and maxi⁡xi≤θ\max_{i}x_{i}\leq\theta.

  2. 2.
    ​

    Each observation satisfies Xi≤θX_{i}\leq\theta with probability one, since the support is [0,θ][0,\theta]. The maximum of numbers, none of which are larger than θ\theta, cannot exceed θ\theta. Thus, θ^=maxi⁡Xi≤θ\hat{\theta}=\max_{i}X_{i}\leq\theta.

  3. 3.
    ​

    Since the distribution of each XiX_{i} is continuous, we have ℙθ⁢(Xi=θ)=0\mathbb{P}_{\theta}(X_{i}=\theta)=0 for all ii. The maximum equals θ\theta only if at least one observation equals this value, and thus we get

    ℙθ⁢(θ^=θ)≤∑i=1nℙθ⁢(Xi=θ)=0.\mathbb{P}_{\theta}(\hat{\theta}=\theta)\leq\sum_{i=1}^{n}\mathbb{P}_{\theta}(% X_{i}=\theta)=0.

    Together with the first part of the solution this shows that θ^<θ\hat{\theta}<\theta with probability one. Thus, the random variable θ−θ^\theta-\hat{\theta} is non-negative and is zero with probability zero. A non-negative random variable with expectation zero equals zero with probability one, and thus we find 𝔼θ⁢(θ−θ^)>0\mathbb{E}_{\theta}(\theta-\hat{\theta})>0, which is the same as 𝔼θ⁢(θ^)<θ\mathbb{E}_{\theta}(\hat{\theta})<\theta. Thus, the estimator never overestimates θ\theta and underestimates θ\theta on average.

∎

C.2 Bias, Mean Squared Error and Consistency

Here we solve the exercises at the end of lecture 2.

Solution to exercise 2.1.
  1. 1.
    ​

    Since 𝔼θ⁢(Xi)=θ\mathbb{E}_{\theta}(X_{i})=\theta, we can use lemma 2.2 to conclude 𝔼θ⁢(X¯)=θ\mathbb{E}_{\theta}(\bar{X})=\theta. Thus we have bias(θ^)=𝔼θ⁢(X¯)−θ=0\mathop{\mathrm{bias}}\nolimits(\hat{\theta})=\mathbb{E}_{\theta}(\bar{X})-% \theta=0 for all θ\theta and the estimator θ^\hat{\theta} is unbiased.

  2. 2.
    ​

    Since the estimator is unbiased, the mean squared error equals the variance. Using lemma 2.2 and the fact that Varθ(Xi)=θ\mathop{\mathrm{Var}}\nolimits_{\theta}(X_{i})=\theta for the Poisson distribution, we find

    MSE(θ^)=Varθ(X¯)=θn.\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})=\mathop{\mathrm{Var}}\nolimits_{% \theta}(\bar{X})=\frac{\theta}{n}.

    Since this converges to zero as n→∞n\to\infty, we can use theorem 2.10 to conclude that the estimator is consistent.

∎

Solution to exercise 2.2.
  1. 1.
    ​

    Using the given density

    𝔼θ⁢(θ^)=∫0θt⋅n⁢tn−1θn⁢dt=nθn⋅θn+1n+1=n⁢θn+1,\mathbb{E}_{\theta}(\hat{\theta})=\int_{0}^{\theta}t\cdot\frac{n\,t^{n-1}}{% \theta^{n}}\,\mathrm{d}t=\frac{n}{\theta^{n}}\cdot\frac{\theta^{n+1}}{n+1}=% \frac{n\theta}{n+1},

    we find the bias to be bias(θ^)=n⁢θ/(n+1)−θ=−θ/(n+1)\mathop{\mathrm{bias}}\nolimits(\hat{\theta})=n\theta/(n+1)-\theta=-\theta/(n+1). Since this converges to zero as n→∞n\to\infty, the estimator is asymptotically unbiased. Since the bias is negative, the maximum systematically underestimates θ\theta, as expected from exercise 1.3.

  2. 2.
    ​

    We have

    𝔼θ⁢(θ^2)=∫0θt2⋅n⁢tn−1θn⁢dt=n⁢θ2n+2,\mathbb{E}_{\theta}(\hat{\theta}^{2})=\int_{0}^{\theta}t^{2}\cdot\frac{n\,t^{n% -1}}{\theta^{n}}\,\mathrm{d}t=\frac{n\theta^{2}}{n+2},

    and thus

    Var(θ^)=n⁢θ2n+2−(n⁢θn+1)2=n⁢θ2(n+2)⁢(n+1)2,\mathop{\mathrm{Var}}\nolimits(\hat{\theta})=\frac{n\theta^{2}}{n+2}-\Bigl{(}% \frac{n\theta}{n+1}\Bigr{)}^{2}=\frac{n\theta^{2}}{(n+2)(n+1)^{2}},

    where we used (n+1)2−n⁢(n+2)=1(n+1)^{2}-n(n+2)=1 in the last step. By theorem 2.7 we find

    MSE(θ^)\displaystyle\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})
    =Var(θ^)+bias(θ^)2\displaystyle=\mathop{\mathrm{Var}}\nolimits(\hat{\theta})+\mathop{\mathrm{% bias}}\nolimits(\hat{\theta})^{2}
    =n⁢θ2(n+2)⁢(n+1)2+θ2(n+1)2\displaystyle=\frac{n\theta^{2}}{(n+2)(n+1)^{2}}+\frac{\theta^{2}}{(n+1)^{2}}
    =2⁢θ2(n+1)⁢(n+2).\displaystyle=\frac{2\theta^{2}}{(n+1)(n+2)}.

    Since this converges to zero as n→∞n\to\infty, we can conclude that θ^\hat{\theta} is consistent by theorem 2.10.

  3. 3.
    ​

    The mean squared error is of order 1/n21/n^{2}, while the mean squared error of the sample mean is of order 1/n1/n. Thus, the maximum converges to the target faster than the average does. The very fast convergence is caused by the support of the uniform model changing as a function of the parameter. This effect, mentioned in example 1.2, will be seen again in lecture 13.

∎

Solution to exercise 2.3.

Since θ^=n/T\hat{\theta}=n/T, we have 𝔼θ⁢(θ^)=n⁢𝔼θ⁢(1/T)\mathbb{E}_{\theta}(\hat{\theta})=n\,\mathbb{E}_{\theta}(1/T) and using the gamma density of TT we get

𝔼θ⁢(1/T)=∫0∞1x⋅θn⁢xn−1⁢e−θ⁢x(n−1)!⁢dx=θn(n−1)!⁢∫0∞xn−2⁢e−θ⁢x⁢dx.\mathbb{E}_{\theta}(1/T)=\int_{0}^{\infty}\frac{1}{x}\cdot\frac{\theta^{n}x^{n% -1}e^{-\theta x}}{(n-1)!}\,\mathrm{d}x=\frac{\theta^{n}}{(n-1)!}\int_{0}^{% \infty}x^{n-2}e^{-\theta x}\,\mathrm{d}x.

For n≥2n\geq 2, the remaining integral can be evaluated as the gamma integral ∫0∞xn−2⁢e−θ⁢x⁢dx=(n−2)!/θn−1\int_{0}^{\infty}x^{n-2}e^{-\theta x}\,\mathrm{d}x=(n-2)!/\theta^{n-1}. Thus we get

𝔼θ⁢(θ^)=n⁢θn(n−1)!⋅(n−2)!θn−1=n⁢θn−1.\mathbb{E}_{\theta}(\hat{\theta})=\frac{n\,\theta^{n}}{(n-1)!}\cdot\frac{(n-2)% !}{\theta^{n-1}}=\frac{n\theta}{n-1}.

The bias is given by

bias(θ^)=n⁢θn−1−θ=θn−1.\mathop{\mathrm{bias}}\nolimits(\hat{\theta})=\frac{n\theta}{n-1}-\theta=\frac% {\theta}{n-1}.

Since this is strictly positive, we see that θ^\hat{\theta} systematically overestimates θ\theta on average and thus is biased. As n→∞n\to\infty, the bias converges to zero and thus θ^\hat{\theta} is asymptotically unbiased. This completes the solution. ∎

C.3 Likelihood and the Method of Moments

Here we solve the exercises at the end of lecture 4.

Solution to exercise 4.1.
  1. 1.
    ​

    Using the Poisson weights from table 1.1, we find the likelihood to be

    L⁢(θ)=∏i=1nθxixi!⁢e−θ=θ∑ixi∏ixi!⁢e−n⁢θ.L(\theta)=\prod_{i=1}^{n}\frac{\theta^{x_{i}}}{x_{i}!}\,e^{-\theta}=\frac{% \theta^{\sum_{i}x_{i}}}{\prod_{i}x_{i}!}\,e^{-n\theta}.

    Taking logarithms we get

    ℓ⁢(θ)=log⁡θ⁢∑i=1nxi−n⁢θ−∑i=1nlog⁡(xi!),\ell(\theta)=\log\theta\sum_{i=1}^{n}x_{i}-n\theta-\sum_{i=1}^{n}\log(x_{i}!),

    as claimed.

  2. 2.
    ​

    Taking derivatives with respect to θ\theta, and noting that the last term does not depend on θ\theta, we find the score function to be

    ℓ′⁢(θ)=1θ⁢∑i=1nxi−n.\ell^{\prime}(\theta)=\frac{1}{\theta}\sum_{i=1}^{n}x_{i}-n.

    Setting this equal to zero and solving gives θ=(1/n)⁢∑ixi=x¯\theta=(1/n)\sum_{i}x_{i}=\bar{x}. Thus, the score function is zero at the sample mean.

  3. 3.
    ​

    For the Poisson distribution we have μ1⁢(θ)=𝔼θ⁢(X1)=θ\mu_{1}(\theta)=\mathbb{E}_{\theta}(X_{1})=\theta. Thus, the only moment equation we have is θ^=X¯\hat{\theta}=\bar{X}: the method of moments estimator is the sample mean. This is the value at which the score function is zero, i.e. for the Poisson distribution the method of moments coincides with the maximum likelihood estimator from lecture 5.

∎

Solution to exercise 4.2.
  1. 1.
    ​

    Since we have μ1⁢(p)=𝔼p⁢(X1)=1/p\mu_{1}(p)=\mathbb{E}_{p}(X_{1})=1/p and the parameter is scalar, we get the moment equation 1/p^=X¯1/\hat{p}=\bar{X}. Solving this equation we find

    p^=1X¯.\hat{p}=\frac{1}{\bar{X}}.
  2. 2.
    ​

    Since each observation satisfies Xi≥1X_{i}\geq 1, we have X¯≥1\bar{X}\geq 1, with equality if and only if all observations equal one. Thus, the interval of p^=1/X¯\hat{p}=1/\bar{X} values where the estimate is possible is (0,1](0,1]. Since for every p∈(0,1)p\in(0,1) the dataset with values in {1,2,…}\{1,2,\dots\} has positive probability, no estimate in this range is incompatible with the data (this is in contrast to the example of the uniform distribution in the text).

∎

Solution to exercise 4.3.
  1. 1.
    ​

    From table A.2 we know that 𝔼θ⁢(X1)=μ\mathbb{E}_{\theta}(X_{1})=\mu and Varθ(X1)=σ2\mathop{\mathrm{Var}}\nolimits_{\theta}(X_{1})=\sigma^{2}. Thus, the first population moment is μ1⁢(θ)=μ\mu_{1}(\theta)=\mu. Rewriting the relation Var(X1)=𝔼⁢(X12)−𝔼⁢(X1)2\mathop{\mathrm{Var}}\nolimits(X_{1})=\mathbb{E}(X_{1}^{2})-\mathbb{E}(X_{1})^% {2} we find

    μ2⁢(θ)=𝔼θ⁢(X12)=Varθ(X1)+𝔼θ⁢(X1)2=σ2+μ2,\mu_{2}(\theta)=\mathbb{E}_{\theta}\bigl{(}X_{1}^{2}\bigr{)}=\mathop{\mathrm{% Var}}\nolimits_{\theta}(X_{1})+\mathbb{E}_{\theta}(X_{1})^{2}=\sigma^{2}+\mu^{% 2},

    as claimed.

  2. 2.
    ​

    Using the moments from the first part, the moment equations from definition 4.6 are

    μ^=m1=X¯andσ^2+μ^2=m2=1n⁢∑i=1nXi2.\hat{\mu}=m_{1}=\bar{X}\qquad\text{and}\qquad\hat{\sigma}^{2}+\hat{\mu}^{2}=m_% {2}=\frac{1}{n}\sum_{i=1}^{n}X_{i}^{2}.

    From the first equation we find μ^=X¯\hat{\mu}=\bar{X} directly. Substituting this value into the second equation we find

    σ^2=1n⁢∑i=1nXi2−X¯2=1n⁢∑i=1n(Xi−X¯)2,\hat{\sigma}^{2}=\frac{1}{n}\sum_{i=1}^{n}X_{i}^{2}-\bar{X}^{2}=\frac{1}{n}% \sum_{i=1}^{n}(X_{i}-\bar{X})^{2},

    where the last equality sign comes from expanding the square (Xi−X¯)2=Xi2−2⁢Xi⁢X¯+X¯2(X_{i}-\bar{X})^{2}=X_{i}^{2}-2X_{i}\bar{X}+\bar{X}^{2} and using the relation ∑iXi=n⁢X¯\sum_{i}X_{i}=n\bar{X}. These are the estimators given in example 4.8 and thus the solution is complete.

∎

Solution to exercise 4.4.
  1. 1.
    ​

    From table A.2 we know that the uniform model has 𝔼θ⁢(X1)=θ/2\mathbb{E}_{\theta}(X_{1})=\theta/2 and Varθ(X1)=θ2/12\mathop{\mathrm{Var}}\nolimits_{\theta}(X_{1})=\theta^{2}/12. Using lemma 2.2 we find 𝔼θ⁢(θ^)=2⁢𝔼θ⁢(X¯)=2⋅θ/2=θ\mathbb{E}_{\theta}(\hat{\theta})=2\,\mathbb{E}_{\theta}(\bar{X})=2\cdot\theta% /2=\theta and thus the estimator θ^\hat{\theta} is unbiased. Since the estimator is unbiased, the mean squared error equals the variance by theorem 2.7, and using lemma 2.2 again we get

    MSE(θ^)=Var(2⁢X¯)=4⋅θ2/12n=θ23⁢n.\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})=\mathop{\mathrm{Var}}\nolimits(2% \bar{X})=4\cdot\frac{\theta^{2}/12}{n}=\frac{\theta^{2}}{3n}.
  2. 2.
    ​

    The maximum has smaller mean squared error, if and only if

    2⁢θ2(n+1)⁢(n+2)<θ23⁢n.\frac{2\theta^{2}}{(n+1)(n+2)}<\frac{\theta^{2}}{3n}.

    This inequality is equivalent to 6⁢n<(n+1)⁢(n+2)6n<(n+1)(n+2), i.e. to 0<n2−3⁢n+2=(n−1)⁢(n−2)0<n^{2}-3n+2=(n-1)(n-2) for all n≥3n\geq 3. (The equality signs hold for n=1n=1 and n=2n=2.) Thus, for three or more observations the biased maximum is better than the unbiased method of moments estimator, and the 1/n21/n^{2} rate of the maximum makes the difference increase quickly with nn. This is another example of the lesson from lecture 2 that unbiasedness is not sufficient: the moment estimator squashes the data into the sample mean, whereas the maximum uses the fact that the support of the distribution moves with the parameter.

∎

Solution to exercise 4.5.
  1. 1.
    ​

    The distribution N⁢(0,θ2)N(0,\theta^{2}) depends on θ\theta only via θ2\theta^{2}. Thus, for different parameter values θ\theta and −θ-\theta, the distribution is the same, e.g. θ1=1\theta_{1}=1 and θ2=−1\theta_{2}=-1 both lead to N⁢(0,1)N(0,1). Thus, the model is not identifiable.

  2. 2.
    ​

    To make the model identifiable, we can restrict ourselves to Θ′=(0,∞)\Theta^{\prime}=(0,\infty): for 0<θ1<θ20<\theta_{1}<\theta_{2} the distributions N⁢(0,θ12)N(0,\theta_{1}^{2}) and N⁢(0,θ22)N(0,\theta_{2}^{2}) have different variances and thus are different. (The other half line (−∞,0)(-\infty,0) would work equally well.)

  3. 3.
    ​

    Since the first moment 𝔼θ⁢(X1)=0\mathbb{E}_{\theta}(X_{1})=0 for all θ\theta, we cannot use this moment to derive any information about the parameter. We can use the second moment instead. Since μ2⁢(θ)=𝔼θ⁢(X12)=θ2\mu_{2}(\theta)=\mathbb{E}_{\theta}(X_{1}^{2})=\theta^{2}, the moment equation is θ^2=(1/n)⁢∑iXi2\hat{\theta}^{2}=(1/n)\sum_{i}X_{i}^{2} and taking the positive root, as required for Θ′=(0,∞)\Theta^{\prime}=(0,\infty), we get

    θ^=1n⁢∑i=1nXi2.\hat{\theta}=\sqrt{\frac{1}{n}\sum_{i=1}^{n}X_{i}^{2}}\,.

    This completes the solution.

∎

C.4 Maximum Likelihood Estimation

Here we solve the exercises at the end of lecture 5.

Solution to exercise 5.1.
  1. 1.
    ​

    Using the geometric weights, the likelihood function is

    L⁢(p)=∏i=1n(1−p)xi−1⁢p=pn⁢(1−p)∑ixi−n.L(p)=\prod_{i=1}^{n}(1-p)^{x_{i}-1}p=p^{n}\,(1-p)^{\sum_{i}x_{i}-n}.

    Taking logarithms we get ℓ⁢(p)=n⁢log⁡p+(∑ixi−n)⁢log⁡(1−p)\ell(p)=n\log p+\bigl{(}\sum_{i}x_{i}-n\bigr{)}\log(1-p), as claimed.

  2. 2.
    ​

    Taking derivatives we find the score as

    ℓ′⁢(p)=np−∑ixi−n1−p.\ell^{\prime}(p)=\frac{n}{p}-\frac{\sum_{i}x_{i}-n}{1-p}.

    Setting this equal to zero and clearing the denominators we find n⁢(1−p)=p⁢(∑ixi−n)n(1-p)=p\bigl{(}\sum_{i}x_{i}-n\bigr{)}, i.e. n=p⁢∑ixin=p\sum_{i}x_{i}. The only solution to this equation is p=n/∑ixi=1/x¯p=n/\sum_{i}x_{i}=1/\bar{x}. Since all observations satisfy xi≥1x_{i}\geq 1, we have ∑ixi≥n\sum_{i}x_{i}\geq n and thus

    ℓ′′⁢(p)=−np2−∑ixi−n(1−p)2<0\ell^{\prime\prime}(p)=-\frac{n}{p^{2}}-\frac{\sum_{i}x_{i}-n}{(1-p)^{2}}<0

    for all p∈(0,1)p\in(0,1). Thus, ℓ\ell is strictly concave and the root is the global maximum. Thus, the MLE is p^=1/X¯\hat{p}=1/\bar{X}. (If all observations equal one, i.e. if ∑ixi=n\sum_{i}x_{i}=n, then ℓ⁢(p)=n⁢log⁡p\ell(p)=n\log p is strictly increasing and no maximiser exists in the open interval (0,1)(0,1). The same boundary phenomenon as for the Bernoulli distribution in the text occurs. In this case, the formula p^=1/X¯\hat{p}=1/\bar{X} gives the boundary value 11.)

  3. 3.
    ​

    In exercise 4.2 we have found the method of moments estimator to be p^=1/X¯\hat{p}=1/\bar{X}. Thus, for the geometric distribution, as for the Bernoulli, Poisson and exponential distributions, the two different construction methods coincide.

∎

Solution to exercise 5.2.
  1. 1.
    ​

    The log-likelihood function is

    ℓ⁢(θ)=−2⁢n⁢log⁡θ+∑i=1nlog⁡xi−1θ⁢∑i=1nxi,\ell(\theta)=-2n\log\theta+\sum_{i=1}^{n}\log x_{i}-\frac{1}{\theta}\sum_{i=1}% ^{n}x_{i},

    and the score function is

    ℓ′⁢(θ)=−2⁢nθ+1θ2⁢∑i=1nxi.\ell^{\prime}(\theta)=-\frac{2n}{\theta}+\frac{1}{\theta^{2}}\sum_{i=1}^{n}x_{% i}.

    Setting the score equal to zero, we find ∑ixi=2⁢n⁢θ\sum_{i}x_{i}=2n\theta and thus the only root is θ=x¯/2\theta=\bar{x}/2. Since the score is positive for θ<x¯/2\theta<\bar{x}/2 and negative for θ>x¯/2\theta>\bar{x}/2, this root is the global maximum and the MLE is θ^=X¯/2\hat{\theta}=\bar{X}/2.

  2. 2.
    ​

    Since the model is described by Gamma⁢(2,1/θ)\text{Gamma}(2,1/\theta) in the rate parametrisation, we can use table A.2 to find 𝔼θ⁢(X1)=2⁢θ\mathbb{E}_{\theta}(X_{1})=2\theta. Using lemma 2.2 we find 𝔼θ⁢(X¯)=2⁢θ\mathbb{E}_{\theta}(\bar{X})=2\theta and thus 𝔼θ⁢(θ^)=(1/2)⋅2⁢θ=θ\mathbb{E}_{\theta}(\hat{\theta})=(1/2)\cdot 2\theta=\theta for all θ\theta. Thus, θ^\hat{\theta} is unbiased.

  3. 3.
    ​

    From table A.2 we find the variance to be Varθ(X1)=2⁢θ2\mathop{\mathrm{Var}}\nolimits_{\theta}(X_{1})=2\theta^{2}. Since the estimator is unbiased, the mean squared error equals the variance and using lemma 2.2 again, we find

    MSE(θ^)=Var(X¯2)=14⋅2⁢θ2n=θ22⁢n.\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})=\mathop{\mathrm{Var}}\nolimits% \Bigl{(}\frac{\bar{X}}{2}\Bigr{)}=\frac{1}{4}\cdot\frac{2\theta^{2}}{n}=\frac{% \theta^{2}}{2n}.

    Since this converges to zero as n→∞n\to\infty, we can conclude that the estimator is consistent by theorem 2.10.

∎

Solution to exercise 5.3.
  1. 1.
    ​

    Since θ^=n/T\hat{\theta}=n/T and since TT has gamma density θn⁢xn−1⁢e−θ⁢x/(n−1)!\theta^{n}x^{n-1}e^{-\theta x}/(n-1)! on (0,∞)(0,\infty), we get the same integral as in exercise 2.3, but with one fewer power of xx, i.e. for n≥3n\geq 3 we have

    𝔼θ⁢(θ^2)\displaystyle\mathbb{E}_{\theta}(\hat{\theta}^{2})
    =n2⁢∫0∞1x2⋅θn⁢xn−1⁢e−θ⁢x(n−1)!⁢dx\displaystyle=n^{2}\int_{0}^{\infty}\frac{1}{x^{2}}\cdot\frac{\theta^{n}x^{n-1% }e^{-\theta x}}{(n-1)!}\,\mathrm{d}x
    =n2⁢θn(n−1)!⋅(n−3)!θn−2=n2(n−1)⁢(n−2)⁢θ2,\displaystyle=\frac{n^{2}\theta^{n}}{(n-1)!}\cdot\frac{(n-3)!}{\theta^{n-2}}=% \frac{n^{2}}{(n-1)(n-2)}\,\theta^{2},

    as claimed.

  2. 2.
    ​

    Expanding the square and using the linearity of the expectation, we find

    MSE(c⁢θ^)=𝔼θ⁢((c⁢θ^−θ)2)=c2⁢𝔼θ⁢(θ^2)−2⁢c⁢θ⁢𝔼θ⁢(θ^)+θ2.\mathop{\mathrm{MSE}}\nolimits(c\hat{\theta})=\mathbb{E}_{\theta}\bigl{(}(c% \hat{\theta}-\theta)^{2}\bigr{)}=c^{2}\,\mathbb{E}_{\theta}(\hat{\theta}^{2})-% 2c\theta\,\mathbb{E}_{\theta}(\hat{\theta})+\theta^{2}.

    In the text we have seen that 𝔼θ⁢(θ^)=n⁢θ/(n−1)\mathbb{E}_{\theta}(\hat{\theta})=n\theta/(n-1), and using this value and the second moment from the previous part of the question, we get

    MSE(c⁢θ^)=(n2(n−1)⁢(n−2)⁢c2−2⁢nn−1⁢c+1)⁢θ2.\mathop{\mathrm{MSE}}\nolimits(c\hat{\theta})=\Bigl{(}\frac{n^{2}}{(n-1)(n-2)}% \,c^{2}-\frac{2n}{n-1}\,c+1\Bigr{)}\theta^{2}.

    For the MLE we set c=1c=1 and thus we get

    MSE(θ^)=(n2(n−1)⁢(n−2)−2⁢nn−1+1)⁢θ2=n+2(n−1)⁢(n−2)⁢θ2.\mathop{\mathrm{MSE}}\nolimits(\hat{\theta})=\Bigl{(}\frac{n^{2}}{(n-1)(n-2)}-% \frac{2n}{n-1}+1\Bigr{)}\theta^{2}=\frac{n+2}{(n-1)(n-2)}\,\theta^{2}.

    For the unbiased estimator we set c=(n−1)/nc=(n-1)/n and thus we get the mean squared error as the variance

    MSE(θ~)=(n−1n−2−2+1)⁢θ2=θ2n−2.\mathop{\mathrm{MSE}}\nolimits(\tilde{\theta})=\Bigl{(}\frac{n-1}{n-2}-2+1% \Bigr{)}\theta^{2}=\frac{\theta^{2}}{n-2}.

    Dividing the two expressions we get MSE(θ~)/MSE(θ^)=(n−1)/(n+2)\mathop{\mathrm{MSE}}\nolimits(\tilde{\theta})/\mathop{\mathrm{MSE}}\nolimits(% \hat{\theta})=(n-1)/(n+2), which is always smaller than 11 for n≥3n\geq 3.

  3. 3.
    ​

    The expression inside the brackets is a quadratic function of cc with positive leading coefficient, so the minimum is at the point where the derivative is zero. This point is

    c=nn−1⋅(n−1)⁢(n−2)n2=n−2n.c=\frac{n}{n-1}\cdot\frac{(n-1)(n-2)}{n^{2}}=\frac{n-2}{n}.

    Substituting this value back into the expression inside the brackets we find

    MSE(n−2n⁢θ^)=(1−n−2n−1)⁢θ2=θ2n−1.\mathop{\mathrm{MSE}}\nolimits\Bigl{(}\frac{n-2}{n}\,\hat{\theta}\Bigr{)}=% \Bigl{(}1-\frac{n-2}{n-1}\Bigr{)}\theta^{2}=\frac{\theta^{2}}{n-1}.

    Since θ2/(n−1)<θ2/(n−2)\theta^{2}/(n-1)<\theta^{2}/(n-2), this value is better than the unbiased estimator (c=(n−1)/nc=(n-1)/n) and the MLE (c=1c=1), as shown in the previous part of the question. Since the winning estimator is smaller than the MLE, we know that the estimator is biased downwards. This is the same effect as we have seen for the normal variance in section 2.3, where the optimal divisor was n+1n+1 instead of the unbiased value n−1n-1 or the value found by maximum likelihood, i.e. nn: the mean squared error can be minimised by a deliberately biased estimator, and the amount of bias depends on the family of distributions under consideration. There is no one-size-fits-all solution for this kind of problem.

∎

Solution to exercise 5.4.
  1. 1.
    ​

    Up to the factor ν+1\nu+1, which we can cancel, the score is a/(ν+a2)+b/(ν+b2)a/(\nu+a^{2})+b/(\nu+b^{2}). Multiplying through by the positive quantity (ν+a2)⁢(ν+b2)(\nu+a^{2})(\nu+b^{2}), the likelihood equation becomes

    a⁢(ν+b2)+b⁢(ν+a2)=ν⁢(a+b)+a⁢b⁢(a+b)=(a+b)⁢(ν+a⁢b)=0,a(\nu+b^{2})+b(\nu+a^{2})=\nu(a+b)+ab(a+b)=(a+b)(\nu+ab)=0,

    as claimed.

  2. 2.
    ​

    For θ=m+t\theta=m+t we have a=d−ta=d-t and b=−d−tb=-d-t, and up to an additive constant the log-likelihood is −ν+12⁢(log⁡(ν+a2)+log⁡(ν+b2))-\frac{\nu+1}{2}\bigl{(}\log(\nu+a^{2})+\log(\nu+b^{2})\bigr{)}. The product inside the logarithms is

    (ν+a2)⁢(ν+b2)\displaystyle(\nu+a^{2})(\nu+b^{2})
    =(ν+d2+t2−2⁢d⁢t)⁢(ν+d2+t2+2⁢d⁢t)\displaystyle=\bigl{(}\nu+d^{2}+t^{2}-2dt\bigr{)}\bigl{(}\nu+d^{2}+t^{2}+2dt% \bigr{)}
    =(ν+d2+t2)2−4⁢d2⁢t2=(ν+d2)2+2⁢(ν−d2)⁢t2+t4,\displaystyle=(\nu+d^{2}+t^{2})^{2}-4d^{2}t^{2}=(\nu+d^{2})^{2}+2(\nu-d^{2})\,% t^{2}+t^{4},

    which is the polynomial P⁢(t)P(t) and thus ℓ⁢(m+t)=C−ν+12⁢log⁡P⁢(t)\ell(m+t)=C-\frac{\nu+1}{2}\log P(t).

  3. 3.
    ​

    Since the logarithm is increasing, ℓ⁢(m+t)\ell(m+t) takes its largest value when P⁢(t)P(t) is smallest. Let u=t2≥0u=t^{2}\geq 0. Then we have P=(ν+d2)2+2⁢(ν−d2)⁢u+u2P=(\nu+d^{2})^{2}+2(\nu-d^{2})u+u^{2}, i.e. we get a quadratic function of uu with minimum at u=d2−νu=d^{2}-\nu. If d2<νd^{2}<\nu, this minimum is at a negative value of uu and thus PP is increasing on u≥0u\geq 0 and takes its minimum at t=0t=0. The midpoint mm is then the only maximiser of ℓ\ell and, since the factor ν+a⁢b=ν−d2+t2\nu+ab=\nu-d^{2}+t^{2} from the first part of the solution has no real zero, mm is the only root of the likelihood equation. If d2>νd^{2}>\nu, then PP takes its minimum at u=d2−νu=d^{2}-\nu, i.e. at t=±d2−νt=\pm\sqrt{d^{2}-\nu} which are the zeros of ν+a⁢b\nu+ab. These two values are the maximum likelihood estimates, and since PP is even in tt, they have equal likelihood. Between these two values, at t=0t=0, the polynomial PP has a local maximum, because the coefficient 2⁢(ν−d2)2(\nu-d^{2}) of t2t^{2} is negative. Thus, ℓ\ell has a local minimum at the midpoint. This completes the proof.

∎

Solution to exercise 5.5.

Since gg is a bijection, there exists an inverse function g−1:Ψ→Θg^{-1}\colon\Psi\to\Theta and we can re-index the model, using ψ\psi instead of θ\theta, i.e. if we let f~⁢(x;ψ)=f⁢(x;g−1⁢(ψ))\tilde{f}(x;\psi)=f\bigl{(}x;g^{-1}(\psi)\bigr{)} the family of distributions is unchanged, but the likelihood function is L~⁢(ψ)=L⁢(g−1⁢(ψ))\tilde{L}(\psi)=L\bigl{(}g^{-1}(\psi)\bigr{)}. For any ψ∈Ψ\psi\in\Psi, the value g−1⁢(ψ)g^{-1}(\psi) falls into the set Θ\Theta and by the definition of θ^\hat{\theta} we have

L~⁢(ψ)=L⁢(g−1⁢(ψ))≤L⁢(θ^)=L~⁢(g⁢(θ^)).\tilde{L}(\psi)=L\bigl{(}g^{-1}(\psi)\bigr{)}\leq L(\hat{\theta})=\tilde{L}% \bigl{(}g(\hat{\theta})\bigr{)}.

Thus, g⁢(θ^)g(\hat{\theta}) is a maximum of the function L~\tilde{L} on the set Ψ\Psi and consequently is a maximum likelihood estimator for ψ\psi. This completes the proof. ∎

Solution to exercise 5.6.
  1. 1.
    ​

    We have s+t≥ss+t\geq s and thus the event {X>s+t}\{X>s+t\} is contained in the event {X>s}\{X>s\}. We get

    ℙθ⁢(X>s+t⁢|X>⁢s)=ℙθ⁢(X>s+t)ℙθ⁢(X>s)=e−θ⁢(s+t)e−θ⁢s=e−θ⁢t=ℙθ⁢(X>t),\mathbb{P}_{\theta}(X>s+t\mskip 1.0mu|\mskip 1.0muX>s)=\frac{\mathbb{P}_{% \theta}(X>s+t)}{\mathbb{P}_{\theta}(X>s)}=\frac{e^{-\theta(s+t)}}{e^{-\theta s% }}=e^{-\theta t}=\mathbb{P}_{\theta}(X>t),

    i.e. the unconditional probability to survive beyond time tt. We can conclude from this that, if a lifetime has already passed time ss, the remaining lifetime has the same distribution as a fresh lifetime does. In other words, the time already elapsed does not change the distribution of the remaining lifetime.

  2. 2.
    ​

    The conditional density of XX given the event {X>c}\{X>c\} is the density restricted to the set (c,∞)(c,\infty), normalised to be a density again. We get f⁢(x;θ)/ℙθ⁢(X>c)f(x;\theta)/\mathbb{P}_{\theta}(X>c) for x>cx>c. Thus we have

    𝔼θ⁢(X⁢|X>⁢c)=1e−θ⁢c⁢∫c∞x⁢θ⁢e−θ⁢x⁢dx.\mathbb{E}_{\theta}(X\mskip 1.0mu|\mskip 1.0muX>c)=\frac{1}{e^{-\theta c}}\int% _{c}^{\infty}x\,\theta e^{-\theta x}\,\mathrm{d}x.

    Using integration by parts, with u=xu=x and v′=θ⁢e−θ⁢xv^{\prime}=\theta e^{-\theta x}, we find v=−e−θ⁢xv=-e^{-\theta x} and thus

    ∫c∞x⁢θ⁢e−θ⁢x⁢dx=[−x⁢e−θ⁢x]c∞+∫c∞e−θ⁢x⁢dx=c⁢e−θ⁢c+e−θ⁢cθ.\int_{c}^{\infty}x\,\theta e^{-\theta x}\,\mathrm{d}x=\Bigl{[}-x\,e^{-\theta x% }\Bigr{]}_{c}^{\infty}+\int_{c}^{\infty}e^{-\theta x}\,\mathrm{d}x=c\,e^{-% \theta c}+\frac{e^{-\theta c}}{\theta}.

    Dividing by e−θ⁢ce^{-\theta c} gives 𝔼θ⁢(X⁢|X>⁢c)=c+1/θ\mathbb{E}_{\theta}(X\mskip 1.0mu|\mskip 1.0muX>c)=c+1/\theta as required. This result is consistent with the first part of the solution: given survival until time cc, the remaining lifetime X−cX-c is distributed like a fresh exponential lifetime. Thus, the conditional expectation is the unconditional mean 1/θ1/\theta, plus the time cc already elapsed, giving c+1/θc+1/\theta. This completes the solution.

∎

Solution to exercise 5.7.
  1. 1.
    ​

    From the solution to exercise 4.1 we know that the score is given by ℓ′⁢(θ)=(1/θ)⁢∑ixi−n\ell^{\prime}(\theta)=(1/\theta)\sum_{i}x_{i}-n and this score equals zero at θ=x¯\theta=\bar{x}. Let S=∑ixi≥1S=\sum_{i}x_{i}\geq 1. Then the second derivative is ℓ′′⁢(θ)=−S/θ2<0\ell^{\prime\prime}(\theta)=-S/\theta^{2}<0 for all θ>0\theta>0 and thus ℓ\ell is strictly concave. This shows that the root is the unique global maximum and thus the MLE is given by θ^=X¯\hat{\theta}=\bar{X}.

  2. 2.
    ​

    The map g⁢(θ)=e−θg(\theta)=e^{-\theta} is continuous and strictly decreasing and therefore maps (0,∞)(0,\infty) onto (0,1)(0,1) bijectively. Thus, by theorem 5.6, the MLE for p0=g⁢(θ)p_{0}=g(\theta) is given by

    p^0=e−θ^=e−X¯.\hat{p}_{0}=e^{-\hat{\theta}}=e^{-\bar{X}}.
  3. 3.
    ​

    Using the independence of the observations and the given identity with z=e−1/nz=e^{-1/n} we get

    𝔼θ⁢(e−X¯)=𝔼θ⁢(∏i=1ne−Xi/n)=∏i=1n𝔼θ⁢((e−1/n)Xi)=exp⁡(n⁢θ⁢(e−1/n−1)).\mathbb{E}_{\theta}\bigl{(}e^{-\bar{X}}\bigr{)}=\mathbb{E}_{\theta}\Bigl{(}% \prod_{i=1}^{n}e^{-X_{i}/n}\Bigr{)}=\prod_{i=1}^{n}\mathbb{E}_{\theta}\Bigl{(}% \bigl{(}e^{-1/n}\bigr{)}^{X_{i}}\Bigr{)}=\exp\bigl{(}n\theta(e^{-1/n}-1)\bigr{% )}.

    Since eu>1+ue^{u}>1+u for all u≠0u\neq 0, we can choose u=−1/nu=-1/n to get e−1/n−1>−1/ne^{-1/n}-1>-1/n and thus n⁢(e−1/n−1)>−1n(e^{-1/n}-1)>-1. This gives

    𝔼θ⁢(p^0)=exp⁡(n⁢θ⁢(e−1/n−1))>e−θ=p0,\mathbb{E}_{\theta}(\hat{p}_{0})=\exp\bigl{(}n\theta(e^{-1/n}-1)\bigr{)}>e^{-% \theta}=p_{0},

    and thus the estimator systematically overestimates p0p_{0} on average and is biased for all nn. On the other hand, expanding the exponential we find n⁢(e−1/n−1)=−1+1/(2⁢n)−⋯→−1n(e^{-1/n}-1)=-1+1/(2n)-\dotsb\to-1 as n→∞n\to\infty and thus 𝔼θ⁢(p^0)→e−θ\mathbb{E}_{\theta}(\hat{p}_{0})\to e^{-\theta} and the estimator is asymptotically unbiased. There is no contradiction to the unbiasedness of X¯\bar{X}: unbiasedness is not preserved under the nonlinear map gg, as discussed after theorem 5.6, whereas the maximum likelihood property is.

∎

Solution to exercise 5.8.
  1. 1.
    ​

    Solving the equation ℙθ⁢(X1≤m)=1−e−θ⁢m=1/2\mathbb{P}_{\theta}(X_{1}\leq m)=1-e^{-\theta m}=1/2 for mm we find e−θ⁢m=1/2e^{-\theta m}=1/2 and taking logarithms we find m=(log⁡2)/θm=(\log 2)/\theta. The map g⁢(θ)=(log⁡2)/θg(\theta)=(\log 2)/\theta is continuous and strictly decreasing on (0,∞)(0,\infty) and since g⁢(θ)→∞g(\theta)\to\infty as θ→0\theta\to 0 and g⁢(θ)→0g(\theta)\to 0 as θ→∞\theta\to\infty, it maps (0,∞)(0,\infty) onto itself in a bijective way.

  2. 2.
    ​

    From example 5.2 we know that the MLE for the rate is given by θ^=1/X¯\hat{\theta}=1/\bar{X}. Since gg is a bijection, theorem 5.6 shows that the MLE for the median m=g⁢(θ)m=g(\theta) is given by

    m^=g⁢(θ^)=log⁡2θ^=(log⁡2)⁢X¯.\hat{m}=g(\hat{\theta})=\frac{\log 2}{\hat{\theta}}=(\log 2)\,\bar{X}.
  3. 3.
    ​

    From lemma 2.2 we know 𝔼θ⁢(X¯)=𝔼θ⁢(X1)=1/θ\mathbb{E}_{\theta}(\bar{X})=\mathbb{E}_{\theta}(X_{1})=1/\theta and thus 𝔼θ⁢(m^)=(log⁡2)/θ=m\mathbb{E}_{\theta}(\hat{m})=(\log 2)/\theta=m for all θ\theta and thus m^\hat{m} is unbiased. There is no contradiction: unbiasedness is a property of a parametrisation and is in general lost under nonlinear transformations of the parameter. Thus, the bias of the MLE θ^=1/X¯\hat{\theta}=1/\bar{X} for the rate is not relevant for the bias of the estimator m^=(log⁡2)⁢X¯\hat{m}=(\log 2)\,\bar{X} for the median, which is a linear function of X¯\bar{X}.

∎

C.5 Exponential Families

Solution to exercise 7.1.
  1. 1.
    ​

    For x∈{1,2,…}x\in\{1,2,\dots\} we can write

    f⁢(x;p)\displaystyle f(x;p)
    =(1−p)x−1⁢p=exp⁡((x−1)⁢log⁡(1−p)+log⁡p)\displaystyle=(1-p)^{x-1}p=\exp\bigl{(}(x-1)\log(1-p)+\log p\bigr{)}
    =exp⁡(x⁢log⁡(1−p)−log⁡1−pp).\displaystyle=\exp\Bigl{(}x\log(1-p)-\log\frac{1-p}{p}\Bigr{)}.

    This is the form of definition 7.1 with h⁢(x)=1h(x)=1 for x∈{1,2,…}x\in\{1,2,\dots\} and h⁢(x)=0h(x)=0 otherwise, natural statistic T⁢(x)=xT(x)=x, natural parameter η⁢(p)=log⁡(1−p)\eta(p)=\log(1-p) and B⁢(p)=log⁡((1−p)/p)B(p)=\log\bigl{(}(1-p)/p\bigr{)}. By lemma 7.5 the natural statistic of a sample is Tn=∑i=1nXiT_{n}=\sum_{i=1}^{n}X_{i}.

  2. 2.
    ​

    For η<0\eta<0 the geometric series gives

    ∑x=1∞eη⁢x=eη1−eη,\sum_{x=1}^{\infty}e^{\eta x}=\frac{e^{\eta}}{1-e^{\eta}},

    but for η≥0\eta\geq 0 the series diverges. Thus we have 𝒩︀=(−∞,0)\mathcal{N}=(-\infty,0) and

    K⁢(η)=η−log⁡(1−eη),η<0.K(\eta)=\eta-\log(1-e^{\eta}),\qquad\eta<0.

    Substituting η⁢(p)=log⁡(1−p)\eta(p)=\log(1-p), so that eη=1−pe^{\eta}=1-p, we get K⁢(η⁢(p))=log⁡(1−p)−log⁡p=B⁢(p)K\bigl{(}\eta(p)\bigr{)}=\log(1-p)-\log p=B(p) as required. As pp ranges over (0,1)(0,1), the natural parameter log⁡(1−p)\log(1-p) ranges over the entire open interval (−∞,0)(-\infty,0), and thus the family has full rank.

  3. 3.
    ​

    We have

    K′⁢(η)=1+eη1−eη=11−eηandK′′⁢(η)=eη(1−eη)2.K^{\prime}(\eta)=1+\frac{e^{\eta}}{1-e^{\eta}}=\frac{1}{1-e^{\eta}}\qquad\text% {and}\qquad K^{\prime\prime}(\eta)=\frac{e^{\eta}}{(1-e^{\eta})^{2}}.

    With eη=1−pe^{\eta}=1-p, by theorem 7.7, we find 𝔼p⁢(X)=1/p\mathbb{E}_{p}(X)=1/p and Varp(X)=(1−p)/p2\mathop{\mathrm{Var}}\nolimits_{p}(X)=(1-p)/p^{2}, the values given in table A.1. The mean 1/p1/p is the moment considered in exercise 4.2.

∎

Solution to exercise 7.2.
  1. 1.
    ​

    For x∈{0,1,…,m}x\in\{0,1,\dots,m\} we can write

    f⁢(x;θ)=(mx)⁢θx⁢(1−θ)m−x=(mx)⁢exp⁡(x⁢log⁡θ1−θ+m⁢log⁡(1−θ)).f(x;\theta)=\binom{m}{x}\theta^{x}(1-\theta)^{m-x}=\binom{m}{x}\exp\Bigl{(}x% \log\frac{\theta}{1-\theta}+m\log(1-\theta)\Bigr{)}.

    This is the form of definition 7.1 with h⁢(x)=(mx)h(x)=\binom{m}{x} for x∈{0,1,…,m}x\in\{0,1,\dots,m\} and h⁢(x)=0h(x)=0 otherwise, natural statistic T⁢(x)=xT(x)=x, natural parameter η⁢(θ)=log⁡(θ/(1−θ))\eta(\theta)=\log\bigl{(}\theta/(1-\theta)\bigr{)}, the same log-odds as in example 7.2, and B⁢(θ)=−m⁢log⁡(1−θ)B(\theta)=-m\log(1-\theta).

  2. 2.
    ​

    Using the binomial theorem, we find

    ∑x=0m(mx)⁢eη⁢x=(1+eη)m\sum_{x=0}^{m}\binom{m}{x}e^{\eta x}=(1+e^{\eta})^{m}

    for all η∈ℝ\eta\in\mathbb{R}. Thus, 𝒩︀=ℝ\mathcal{N}=\mathbb{R} and K⁢(η)=m⁢log⁡(1+eη)K(\eta)=m\log(1+e^{\eta}). This is mm times the cumulant function of the Bernoulli distribution.

  3. 3.
    ​

    Taking derivatives, we get

    K′⁢(η)=m⁢eη1+eηandK′′⁢(η)=m⁢eη(1+eη)2.K^{\prime}(\eta)=\frac{me^{\eta}}{1+e^{\eta}}\qquad\text{and}\qquad K^{\prime% \prime}(\eta)=\frac{me^{\eta}}{(1+e^{\eta})^{2}}.

    Since eη/(1+eη)=θe^{\eta}/(1+e^{\eta})=\theta and 1/(1+eη)=1−θ1/(1+e^{\eta})=1-\theta, theorem 7.7 gives 𝔼θ⁢(X)=m⁢θ\mathbb{E}_{\theta}(X)=m\theta and Varθ(X)=m⁢θ⁢(1−θ)\mathop{\mathrm{Var}}\nolimits_{\theta}(X)=m\theta(1-\theta), in agreement with table A.1. For a sample of size nn the natural statistic is Tn=∑iXiT_{n}=\sum_{i}X_{i} and, by the remark after theorem 7.7, we have 𝔼θ⁢(Tn)=n⁢m⁢θ\mathbb{E}_{\theta}(T_{n})=nm\theta and Varθ(Tn)=n⁢m⁢θ⁢(1−θ)\mathop{\mathrm{Var}}\nolimits_{\theta}(T_{n})=nm\theta(1-\theta).

∎

Solution to exercise 7.3.
  1. 1.
    ​

    We can write xα−1=e(α−1)⁢log⁡xx^{\alpha-1}=e^{(\alpha-1)\log x} and, collecting all factors into one exponential, for x>0x>0 we find

    f⁢(x;α,β)\displaystyle f(x;\alpha,\beta)
    =βαΓ⁢(α)⁢e(α−1)⁢log⁡x⁢e−β⁢x\displaystyle=\frac{\beta^{\alpha}}{\Gamma(\alpha)}\,e^{(\alpha-1)\log x}\,e^{% -\beta x}
    =exp⁡((α−1)⁢log⁡x+(−β)⁢x−(log⁡Γ⁢(α)−α⁢log⁡β)).\displaystyle=\exp\Bigl{(}(\alpha-1)\log x+(-\beta)\,x-\bigl{(}\log\Gamma(% \alpha)-\alpha\log\beta\bigr{)}\Bigr{)}.

    This is the form of definition 7.8 with k=2k=2, with h⁢(x)=1h(x)=1 for x>0x>0 and h⁢(x)=0h(x)=0 otherwise, natural statistic (T1⁢(x),T2⁢(x))=(log⁡x,x)\bigl{(}T_{1}(x),T_{2}(x)\bigr{)}=(\log x,\,x), natural parameter (η1⁢(α,β),η2⁢(α,β))=(α−1,−β)\bigl{(}\eta_{1}(\alpha,\beta),\eta_{2}(\alpha,\beta)\bigr{)}=(\alpha-1,\,-\beta) and B⁢(α,β)=log⁡Γ⁢(α)−α⁢log⁡βB(\alpha,\beta)=\log\Gamma(\alpha)-\alpha\log\beta.

  2. 2.
    ​

    As α\alpha varies over (0,∞)(0,\infty), the first coordinate η1=α−1\eta_{1}=\alpha-1 varies over (−1,∞)(-1,\infty) and as β\beta varies over (0,∞)(0,\infty), the second coordinate η2=−β\eta_{2}=-\beta varies over (−∞,0)(-\infty,0), independently of each other. Thus, the set of natural parameter values is the open quadrant {(η1,η2)∣η1>−1,η2<0}\{(\eta_{1},\eta_{2})\mid\eta_{1}>-1,\ \eta_{2}<0\}. This is a non-empty open subset of ℝ2\mathbb{R}^{2} and the family has full rank as in definition 7.10.

  3. 3.
    ​

    Knowing α\alpha, we can include the factor xα−1x^{\alpha-1} into the function hh: for x>0x>0 we have

    f⁢(x;β)=xα−1⁢exp⁡((−β)⁢x−(log⁡Γ⁢(α)−α⁢log⁡β)),f(x;\beta)=x^{\alpha-1}\exp\Bigl{(}(-\beta)\,x-\bigl{(}\log\Gamma(\alpha)-% \alpha\log\beta\bigr{)}\Bigr{)},

    This is a one-parameter exponential family in the sense of definition 7.1, with h⁢(x)=xα−1h(x)=x^{\alpha-1} for x>0x>0, natural statistic T⁢(x)=xT(x)=x and natural parameter −β-\beta.

∎

Solution to exercise 7.4.
  1. 1.
    ​

    Using the formula K⁢(η)=−log⁡(−η)K(\eta)=-\log(-\eta) on 𝒩︀=(−∞,0)\mathcal{N}=(-\infty,0) we find K′⁢(η)=−1/ηK^{\prime}(\eta)=-1/\eta and K′′⁢(η)=1/η2K^{\prime\prime}(\eta)=1/\eta^{2}. Substituting the natural parameter η=−θ\eta=-\theta from example 7.4, we can use theorem 7.7 to find 𝔼θ⁢(X)=1/θ\mathbb{E}_{\theta}(X)=1/\theta and Varθ(X)=1/θ2\mathop{\mathrm{Var}}\nolimits_{\theta}(X)=1/\theta^{2}. This is the mean and variance of the exponential distribution with rate θ\theta.

  2. 2.
    ​

    For h⁢(x)=(2⁢π⁢σ2)−1/2⁢e−x2/(2⁢σ2)h(x)=(2\pi\sigma^{2})^{-1/2}e^{-x^{2}/(2\sigma^{2})} and T⁢(x)=xT(x)=x we have to evaluate

    Z⁢(η)=∫−∞∞12⁢π⁢σ2⁢exp⁡(−x2−2⁢σ2⁢η⁢x2⁢σ2)⁢dx.Z(\eta)=\int_{-\infty}^{\infty}\frac{1}{\sqrt{2\pi\sigma^{2}}}\exp\Bigl{(}-% \frac{x^{2}-2\sigma^{2}\eta x}{2\sigma^{2}}\Bigr{)}\,\mathrm{d}x.

    We can complete the square to find x2−2⁢σ2⁢η⁢x=(x−σ2⁢η)2−σ4⁢η2x^{2}-2\sigma^{2}\eta x=(x-\sigma^{2}\eta)^{2}-\sigma^{4}\eta^{2} and thus

    Z⁢(η)=exp⁡(σ2⁢η22)⁢∫−∞∞12⁢π⁢σ2⁢exp⁡(−(x−σ2⁢η)22⁢σ2)⁢dx=exp⁡(σ2⁢η22),Z(\eta)=\exp\Bigl{(}\frac{\sigma^{2}\eta^{2}}{2}\Bigr{)}\int_{-\infty}^{\infty% }\frac{1}{\sqrt{2\pi\sigma^{2}}}\exp\Bigl{(}-\frac{(x-\sigma^{2}\eta)^{2}}{2% \sigma^{2}}\Bigr{)}\,\mathrm{d}x=\exp\Bigl{(}\frac{\sigma^{2}\eta^{2}}{2}\Bigr% {)},

    since the remaining integrand is the density of N⁢(σ2⁢η,σ2)N(\sigma^{2}\eta,\sigma^{2}) and thus integrates to one. The integral is finite for all η\eta and thus 𝒩︀=ℝ\mathcal{N}=\mathbb{R} and K⁢(η)=log⁡Z⁢(η)=σ2⁢η2/2K(\eta)=\log Z(\eta)=\sigma^{2}\eta^{2}/2. Taking derivatives we find K′⁢(η)=σ2⁢ηK^{\prime}(\eta)=\sigma^{2}\eta and K′′⁢(η)=σ2K^{\prime\prime}(\eta)=\sigma^{2} and substituting η=μ/σ2\eta=\mu/\sigma^{2} in theorem 7.7 we get 𝔼μ⁢(X)=μ\mathbb{E}_{\mu}(X)=\mu and Varμ(X)=σ2\mathop{\mathrm{Var}}\nolimits_{\mu}(X)=\sigma^{2}, as expected. Finally, K⁢(η⁢(μ))=σ2⁢(μ/σ2)2/2=μ2/(2⁢σ2)K\bigl{(}\eta(\mu)\bigr{)}=\sigma^{2}(\mu/\sigma^{2})^{2}/2=\mu^{2}/(2\sigma^{% 2}), which is the function B⁢(μ)B(\mu) from example 7.9.

∎

Solution to exercise 7.5.
  1. 1.
    ​

    By lemma 7.5, applied in the natural parametrisation where B=KB=K, the log-likelihood is

    ℓ⁢(η)=∑i=1nlog⁡h⁢(xi)+η⁢Tn−n⁢K⁢(η),Tn=∑i=1nT⁢(xi),\ell(\eta)=\sum_{i=1}^{n}\log h(x_{i})+\eta\,T_{n}-nK(\eta),\qquad T_{n}=\sum_% {i=1}^{n}T(x_{i}),

    and thus the score is ℓ′⁢(η)=Tn−n⁢K′⁢(η)\ell^{\prime}(\eta)=T_{n}-nK^{\prime}(\eta). The likelihood equation ℓ′⁢(η)=0\ell^{\prime}(\eta)=0 is therefore K′⁢(η)=Tn/nK^{\prime}(\eta)=T_{n}/n, and by theorem 7.7 the left-hand side equals 𝔼η⁢(T⁢(X1))\mathbb{E}_{\eta}\bigl{(}T(X_{1})\bigr{)}. This is the stated equation.

  2. 2.
    ​

    Differentiating once more and using theorem 7.7 again, we find

    ℓ′′⁢(η)=−n⁢K′′⁢(η)=−n⁢Varη(T⁢(X1))<0\ell^{\prime\prime}(\eta)=-nK^{\prime\prime}(\eta)=-n\mathop{\mathrm{Var}}% \nolimits_{\eta}\bigl{(}T(X_{1})\bigr{)}<0

    for all η∈𝒩︀\eta\in\mathcal{N}. Thus ℓ\ell is strictly concave on the interval 𝒩︀\mathcal{N}, and a strictly concave function has at most one critical point, which is then its global maximum. Any root of the likelihood equation is therefore the unique maximiser of ℓ\ell, i.e. the unique MLE for η\eta.

  3. 3.
    ​

    For the Poisson distribution we have T⁢(x)=xT(x)=x and K′⁢(η)=eηK^{\prime}(\eta)=e^{\eta}, and thus the likelihood equation reads eη=x¯e^{\eta}=\bar{x}, with root η^=log⁡x¯\hat{\eta}=\log\bar{x} whenever x¯>0\bar{x}>0. Since θ=eη\theta=e^{\eta} is a bijection between 𝒩︀=ℝ\mathcal{N}=\mathbb{R} and Θ=(0,∞)\Theta=(0,\infty), theorem 5.6 gives θ^=eη^=X¯\hat{\theta}=e^{\hat{\eta}}=\bar{X}, the MLE of exercise 5.7. For the Bernoulli distribution we have T⁢(x)=xT(x)=x and K′⁢(η)=eη/(1+eη)K^{\prime}(\eta)=e^{\eta}/(1+e^{\eta}), and thus the likelihood equation reads eη/(1+eη)=x¯e^{\eta}/(1+e^{\eta})=\bar{x}, with root η^=log⁡(x¯/(1−x¯))\hat{\eta}=\log\bigl{(}\bar{x}/(1-\bar{x})\bigr{)} whenever 0<x¯<10<\bar{x}<1. Since θ=eη/(1+eη)\theta=e^{\eta}/(1+e^{\eta}) is a bijection between ℝ\mathbb{R} and (0,1)(0,1), theorem 5.6 gives θ^=X¯\hat{\theta}=\bar{X}, the MLE of example 5.3. The restriction 0<x¯<10<\bar{x}<1 is the condition 0<s<n0<s<n discussed after that example: when it fails the likelihood equation has no root and the MLE does not exist in the open parameter space.

  4. 4.
    ​

    The equation of the first part says that the MLE is the parameter value under which the population mean of the natural statistic TT agrees with its sample mean, i.e. the MLE is the method of moments estimator based on the moment 𝔼θ⁢(T⁢(X1))\mathbb{E}_{\theta}\bigl{(}T(X_{1})\bigr{)} rather than on 𝔼θ⁢(X1)\mathbb{E}_{\theta}(X_{1}). Whenever T⁢(x)=xT(x)=x, as for the Bernoulli, Poisson and exponential distributions and for the normal mean with known variance, the two moments are the same, and the MLE coincides with the method of moments estimator from definition 4.6; this is the coincidence observed in the remark at the end of the method of moments section of lecture 4. For the gamma distribution of exercise 7.3 the natural statistic is (log⁡x,x)(\log x,\,x), and thus maximum likelihood matches the means of log⁡X1\log X_{1} and X1X_{1}, whereas the method of moments in example 4.9 matched the means of X1X_{1} and X12X_{1}^{2}. Different moments are matched, and thus the two estimators differ.

∎

Solution to exercise 7.6.
  1. 1.
    ​

    Writing the expectation as an integral (or a sum, for a discrete model), we find

    𝔼η⁢(es⁢T⁢(X))\displaystyle\mathbb{E}_{\eta}\bigl{(}e^{sT(X)}\bigr{)}
    =∫es⁢T⁢(x)⁢h⁢(x)⁢eη⁢T⁢(x)−K⁢(η)⁢dx\displaystyle=\int e^{sT(x)}\,h(x)\,e^{\eta T(x)-K(\eta)}\,\mathrm{d}x
    =e−K⁢(η)⁢∫h⁢(x)⁢e(η+s)⁢T⁢(x)⁢dx=eK⁢(η+s)−K⁢(η),\displaystyle=e^{-K(\eta)}\int h(x)\,e^{(\eta+s)T(x)}\,\mathrm{d}x=e^{K(\eta+s% )-K(\eta)},

    where the last step uses the definition of KK at the point η+s∈𝒩︀\eta+s\in\mathcal{N}. Taking logarithms gives log⁡𝔼η⁢(es⁢T⁢(X))=K⁢(η+s)−K⁢(η)=M⁢(s)\log\mathbb{E}_{\eta}(e^{sT(X)})=K(\eta+s)-K(\eta)=M(s), as claimed.

  2. 2.
    ​

    By the first part we have eM⁢(s)=𝔼η⁢(es⁢T⁢(X))e^{M(s)}=\mathbb{E}_{\eta}(e^{sT(X)}). Differentiating both sides with respect to ss, and taking the derivative inside the expectation on the right, we get

    M′⁢(s)⁢eM⁢(s)\displaystyle M^{\prime}(s)\,e^{M(s)}
    =𝔼η⁢(T⁢(X)⁢es⁢T⁢(X))and\displaystyle=\mathbb{E}_{\eta}\bigl{(}T(X)\,e^{sT(X)}\bigr{)}\qquad\text{and}
    (M′′⁢(s)+M′⁢(s)2)⁢eM⁢(s)\displaystyle\bigl{(}M^{\prime\prime}(s)+M^{\prime}(s)^{2}\bigr{)}e^{M(s)}
    =𝔼η⁢(T⁢(X)2⁢es⁢T⁢(X)).\displaystyle=\mathbb{E}_{\eta}\bigl{(}T(X)^{2}\,e^{sT(X)}\bigr{)}.

    At s=0s=0 we have M⁢(0)=0M(0)=0, and thus M′⁢(0)=𝔼η⁢(T⁢(X))M^{\prime}(0)=\mathbb{E}_{\eta}(T(X)) and M′′⁢(0)=𝔼η⁢(T⁢(X)2)−𝔼η⁢(T⁢(X))2=Varη(T⁢(X))M^{\prime\prime}(0)=\mathbb{E}_{\eta}(T(X)^{2})-\mathbb{E}_{\eta}(T(X))^{2}=% \mathop{\mathrm{Var}}\nolimits_{\eta}(T(X)). On the other hand, from M⁢(s)=K⁢(η+s)−K⁢(η)M(s)=K(\eta+s)-K(\eta) we get directly M′⁢(0)=K′⁢(η)M^{\prime}(0)=K^{\prime}(\eta) and M′′⁢(0)=K′′⁢(η)M^{\prime\prime}(0)=K^{\prime\prime}(\eta). Comparing the two computations gives 𝔼η⁢(T⁢(X))=K′⁢(η)\mathbb{E}_{\eta}(T(X))=K^{\prime}(\eta) and Varη(T⁢(X))=K′′⁢(η)\mathop{\mathrm{Var}}\nolimits_{\eta}(T(X))=K^{\prime\prime}(\eta), which is theorem 7.7.

  3. 3.
    ​

    For the Poisson distribution we have T⁢(x)=xT(x)=x, η=log⁡θ\eta=\log\theta and K⁢(η)=eηK(\eta)=e^{\eta} by example 7.3 and the computation following definition 7.6, and 𝒩︀=ℝ\mathcal{N}=\mathbb{R}, so that η+s∈𝒩︀\eta+s\in\mathcal{N} for every ss. Writing z=esz=e^{s}, the first part gives

    𝔼θ⁢(zX)=𝔼η⁢(es⁢X)=exp⁡(eη+s−eη)=exp⁡(θ⁢z−θ)=eθ⁢(z−1),\mathbb{E}_{\theta}\bigl{(}z^{X}\bigr{)}=\mathbb{E}_{\eta}\bigl{(}e^{sX}\bigr{% )}=\exp\bigl{(}e^{\eta+s}-e^{\eta}\bigr{)}=\exp\bigl{(}\theta z-\theta\bigr{)}% =e^{\theta(z-1)},

    which is the identity used in exercise 5.7.

∎

Solution to exercise 7.7.
  1. 1.
    ​

    Using the notation μ=θ\mu=\theta and σ2=θ2\sigma^{2}=\theta^{2} from example 7.9, we see that the density of N⁢(θ,θ2)N(\theta,\theta^{2}) is of the form given in definition 7.8 with k=2k=2, where h⁢(x)=1h(x)=1, the natural statistic is (x,x2)(x,x^{2}) and the natural parameter is

    (μσ2,−12⁢σ2)=(1θ,−12⁢θ2).\Bigl{(}\frac{\mu}{\sigma^{2}},\,-\frac{1}{2\sigma^{2}}\Bigr{)}=\Bigl{(}\frac{% 1}{\theta},\,-\frac{1}{2\theta^{2}}\Bigr{)}.
  2. 2.
    ​

    Let η1=1/θ\eta_{1}=1/\theta and η2=−1/(2⁢θ2)\eta_{2}=-1/(2\theta^{2}). As θ\theta varies over the interval (0,∞)(0,\infty), the first coordinate η1\eta_{1} also varies over the interval (0,∞)(0,\infty). Substituting θ=1/η1\theta=1/\eta_{1} for the second coordinate, we find

    η2=−12⁢θ2=−η122.\eta_{2}=-\frac{1}{2\theta^{2}}=-\frac{\eta_{1}^{2}}{2}.

    Thus, the set of natural parameter values is the curve {(η1,−η12/2)|η1>0}\bigl{\{}(\eta_{1},-\eta_{1}^{2}/2)\mathrel{\big{|}}\eta_{1}>0\bigr{\}}. Any open subset of ℝ2\mathbb{R}^{2} which contains a point on the curve also contains a small disc around this point, and thus points off the curve, obtained by slightly changing η2\eta_{2} while keeping η1\eta_{1} fixed. Thus the curve does not contain any non-empty open subsets of ℝ2\mathbb{R}^{2}. By definition 7.10, the family is not of full rank, as claimed.

∎

C.6 Sufficiency

Here we solve the exercises at the end of lecture 8.

Solution to exercise 8.1.
  1. 1.
    ​

    If ∑ixi=t\sum_{i}x_{i}=t, the event {X=x}\{X=x\} is contained in the event {T=t}\{T=t\} and we have

    ℙθ⁢(X=x|T=t)=ℙθ⁢(X=x)ℙθ⁢(T=t)=θt⁢(1−θ)n−t(nt)⁢θt⁢(1−θ)n−t=1(nt).\mathbb{P}_{\theta}(X=x\mskip 1.0mu|\mskip 1.0muT=t)=\frac{\mathbb{P}_{\theta}% (X=x)}{\mathbb{P}_{\theta}(T=t)}=\frac{\theta^{t}(1-\theta)^{n-t}}{\binom{n}{t% }\theta^{t}(1-\theta)^{n-t}}=\frac{1}{\binom{n}{t}}.
  2. 2.
    ​

    If ∑ixi≠t\sum_{i}x_{i}\neq t, the events {X=x}\{X=x\} and {T=t}\{T=t\} are disjoint and the conditional probability is zero. Thus the result does not depend on θ\theta and TT is sufficient for θ\theta by definition 8.1.

  3. 3.
    ​

    The joint probability weights are given by

    f⁢(x;θ)=∏i=1nθxi⁢(1−θ)1−xi=θ∑ixi⁢(1−θ)n−∑ixif(x;\theta)=\prod_{i=1}^{n}\theta^{x_{i}}(1-\theta)^{1-x_{i}}=\theta^{\sum_{i}% x_{i}}(1-\theta)^{n-\sum_{i}x_{i}}

    for x∈{0,1}nx\in\{0,1\}^{n}. This function is of the form (8.1) with g⁢(t;θ)=θt⁢(1−θ)n−tg(t;\theta)=\theta^{t}(1-\theta)^{n-t} and h⁢(x)=1h(x)=1. By theorem 8.2 the statistic TT is sufficient, which confirms our direct computation.

∎

Solution to exercise 8.2.
  1. 1.
    ​

    The density of a single observation is f⁢(xi;θ)=θ−1⁢𝟏{0≤xi≤θ}f(x_{i};\theta)=\theta^{-1}\mathbf{1}_{\{0\leq x_{i}\leq\theta\}}. Since the joint density is the product of these densities, and since a product of indicators can be written as the indicator of the intersection of the corresponding events, we find that all nn conditions 0≤xi≤θ0\leq x_{i}\leq\theta are satisfied if and only if the smallest observation is greater than or equal to 0 and the largest observation is less than or equal to θ\theta. Thus we have

    f⁢(x;θ)=θ−n⁢∏i=1n𝟏{0≤xi≤θ}=θ−n⁢ 1{maxi⁡xi≤θ}⁢ 1{mini⁡xi≥0}.f(x;\theta)=\theta^{-n}\prod_{i=1}^{n}\mathbf{1}_{\{0\leq x_{i}\leq\theta\}}=% \theta^{-n}\,\mathbf{1}_{\{\max_{i}x_{i}\leq\theta\}}\,\mathbf{1}_{\{\min_{i}x% _{i}\geq 0\}}.
  2. 2.
    ​

    Let g⁢(t;θ)=θ−n⁢𝟏{t≤θ}g(t;\theta)=\theta^{-n}\mathbf{1}_{\{t\leq\theta\}} and h⁢(x)=𝟏{mini⁡xi≥0}h(x)=\mathbf{1}_{\{\min_{i}x_{i}\geq 0\}}. Then gg depends on the data only through t=maxi⁡xit=\max_{i}x_{i}, the function hh does not depend on θ\theta and from the first part of the solution we know that f⁢(x;θ)=g⁢(maxi⁡xi;θ)⁢h⁢(x)f(x;\theta)=g(\max_{i}x_{i};\theta)\,h(x) for all xx and θ\theta. By theorem 8.2 the statistic X(n)=maxi⁡XiX_{(n)}=\max_{i}X_{i} is sufficient for θ\theta.

  3. 3.
    ​

    The factor 𝟏{mini⁡xi≥0}\mathbf{1}_{\{\min_{i}x_{i}\geq 0\}} depends on the data but does not depend on θ\theta and thus can be included in hh. The factor 𝟏{maxi⁡xi≤θ}\mathbf{1}_{\{\max_{i}x_{i}\leq\theta\}} depends on θ\theta and thus a function involving this factor cannot be in hh. This factor must be included in gg and, since it only depends on the data through maxi⁡xi\max_{i}x_{i}, the maximum is the sufficient statistic resulting from this factorisation. It is a common mistake to write the joint density as θ−n\theta^{-n} and to conclude that no additional statistic is needed, since the indicator is where the data and the parameter interact.

∎

Solution to exercise 8.3.
  1. 1.
    ​

    For x1,…,xn≥0x_{1},\dots,x_{n}\geq 0 the joint density is

    f⁢(x;θ)=∏i=1nθ⁢e−θ⁢xi=θn⁢exp⁡(−θ⁢∑i=1nxi),f(x;\theta)=\prod_{i=1}^{n}\theta e^{-\theta x_{i}}=\theta^{n}\exp\Bigl{(}-% \theta\sum_{i=1}^{n}x_{i}\Bigr{)},

    and it is zero if any xix_{i} is negative. Thus the factorisation (8.1) holds with g⁢(t;θ)=θn⁢e−θ⁢tg(t;\theta)=\theta^{n}e^{-\theta t} and h⁢(x)=𝟏{mini⁡xi≥0}h(x)=\mathbf{1}_{\{\min_{i}x_{i}\geq 0\}}, and T=∑iXiT=\sum_{i}X_{i} is sufficient by theorem 8.2.

  2. 2.
    ​

    For two data vectors xx and yy with non-negative entries, the first part gives

    f⁢(x;θ)f⁢(y;θ)=exp⁡(−θ⁢(∑ixi−∑iyi)).\frac{f(x;\theta)}{f(y;\theta)}=\exp\Bigl{(}-\theta\Bigl{(}\sum_{i}x_{i}-\sum_% {i}y_{i}\Bigr{)}\Bigr{)}.

    If ∑ixi=∑iyi\sum_{i}x_{i}=\sum_{i}y_{i}, the ratio equals 11 for every θ\theta. If the sums differ by d≠0d\neq 0, the ratio is e−θ⁢de^{-\theta d}, which is a strictly monotone function of θ\theta and thus not constant on (0,∞)(0,\infty). The ratio is therefore free of θ\theta if and only if T⁢(x)=T⁢(y)T(x)=T(y) and by theorem 8.6 we find that TT is minimal sufficient.

  3. 3.
    ​

    Since X¯=T/n\bar{X}=T/n, the MLE is given by θ^=1/X¯=n/T\hat{\theta}=1/\bar{X}=n/T, a function of TT. For fixed data, the likelihood function is L⁢(θ)=g⁢(T;θ)⁢h⁢(x)=θn⁢e−θ⁢T⁢h⁢(x)L(\theta)=g(T;\theta)\,h(x)=\theta^{n}e^{-\theta T}\,h(x), where h⁢(x)h(x) is a positive constant not depending on θ\theta. Thus, the maximiser of L⁢(θ)L(\theta) over θ\theta is the same as the maximiser of θn⁢e−θ⁢T\theta^{n}e^{-\theta T}, which depends on the data only through TT. Whatever the MLE is, it can be written as a function of the sufficient statistic by the factorisation in theorem 8.2.

∎

Solution to exercise 8.4.
  1. 1.
    ​

    With μ\mu known, the joint density

    f⁢(x;σ2)=(2⁢π⁢σ2)−n/2⁢exp⁡(−12⁢σ2⁢∑i=1n(xi−μ)2)f(x;\sigma^{2})=(2\pi\sigma^{2})^{-n/2}\exp\Bigl{(}-\frac{1}{2\sigma^{2}}\sum_% {i=1}^{n}(x_{i}-\mu)^{2}\Bigr{)}

    depends on the data only through t=∑i(xi−μ)2t=\sum_{i}(x_{i}-\mu)^{2}, so that (8.1) holds with g⁢(t;σ2)=(2⁢π⁢σ2)−n/2⁢exp⁡(−t/(2⁢σ2))g(t;\sigma^{2})=(2\pi\sigma^{2})^{-n/2}\exp\bigl{(}-t/(2\sigma^{2})\bigr{)} and h⁢(x)=1h(x)=1. By theorem 8.2 the statistic ∑i(Xi−μ)2\sum_{i}(X_{i}-\mu)^{2} is sufficient for σ2\sigma^{2}.

  2. 2.
    ​

    Expanding the squares in the exponent, the joint density is

    f⁢(x;μ,σ2)\displaystyle f(x;\mu,\sigma^{2})
    =(2⁢π⁢σ2)−n/2⁢exp⁡(−12⁢σ2⁢∑i=1n(xi−μ)2)\displaystyle=(2\pi\sigma^{2})^{-n/2}\exp\Bigl{(}-\frac{1}{2\sigma^{2}}\sum_{i% =1}^{n}(x_{i}-\mu)^{2}\Bigr{)}
    =(2⁢π⁢σ2)−n/2⁢exp⁡(μσ2⁢∑ixi−12⁢σ2⁢∑ixi2−n⁢μ22⁢σ2),\displaystyle=(2\pi\sigma^{2})^{-n/2}\exp\Bigl{(}\frac{\mu}{\sigma^{2}}\sum_{i% }x_{i}-\frac{1}{2\sigma^{2}}\sum_{i}x_{i}^{2}-\frac{n\mu^{2}}{2\sigma^{2}}% \Bigr{)},

    which depends on the data only through the pair T=(∑ixi,∑ixi2)T=\bigl{(}\sum_{i}x_{i},\sum_{i}x_{i}^{2}\bigr{)}, and thus the factorisation (8.1) holds with h⁢(x)=1h(x)=1 and with gg equal to the whole expression. By theorem 8.2 the pair TT is sufficient for (μ,σ2)(\mu,\sigma^{2}).

  3. 3.
    ​

    Using the expanded form of the joint density from the previous part, and writing a=∑ixi−∑iyia=\sum_{i}x_{i}-\sum_{i}y_{i} and b=∑ixi2−∑iyi2b=\sum_{i}x_{i}^{2}-\sum_{i}y_{i}^{2}, the ratio of the densities of two data vectors is

    f⁢(x;μ,σ2)f⁢(y;μ,σ2)=exp⁡(a⁢μ−b/2σ2),\frac{f(x;\mu,\sigma^{2})}{f(y;\mu,\sigma^{2})}=\exp\Bigl{(}\frac{a\mu-b/2}{% \sigma^{2}}\Bigr{)},

    since the factors (2⁢π⁢σ2)−n/2(2\pi\sigma^{2})^{-n/2} and exp⁡(−n⁢μ2/(2⁢σ2))\exp\bigl{(}-n\mu^{2}/(2\sigma^{2})\bigr{)} cancel. If a=b=0a=b=0, which is the statement T⁢(x)=T⁢(y)T(x)=T(y), the ratio equals 11 for all parameter values. Conversely, suppose the ratio does not depend on (μ,σ2)(\mu,\sigma^{2}), so that the exponent (a⁢μ−b/2)/σ2(a\mu-b/2)/\sigma^{2} equals a constant cc for all μ∈ℝ\mu\in\mathbb{R} and all σ2>0\sigma^{2}>0. Taking σ2=1\sigma^{2}=1, the expression a⁢μ−b/2a\mu-b/2 is constant in μ\mu, which forces a=0a=0 and then b=−2⁢cb=-2c. With a=0a=0 the exponent is −b/(2⁢σ2)-b/(2\sigma^{2}), and this is constant in σ2\sigma^{2} only if b=0b=0. Thus the ratio is free of the parameters if and only if T⁢(x)=T⁢(y)T(x)=T(y), and TT is minimal sufficient by theorem 8.6.

  4. 4.
    ​

    No. By the previous part TT is minimal sufficient, and thus, by definition 8.5, TT would have to be a function of X¯\bar{X} if X¯\bar{X} were sufficient. It is not: for n=2n=2 the data sets x=(1,−1)x=(1,-1) and y=(2,−2)y=(2,-2) have the same sample mean 0, but ∑ixi2=2\sum_{i}x_{i}^{2}=2 and ∑iyi2=8\sum_{i}y_{i}^{2}=8, so that T⁢(x)≠T⁢(y)T(x)\neq T(y). Two data sets with the same value of X¯\bar{X} can therefore have non-proportional likelihood functions, and X¯\bar{X} alone loses information about σ2\sigma^{2}.

∎

Solution to exercise 8.5.
  1. 1.
    ​

    The exponential distribution with rate θ\theta has distribution function F⁢(y)=1−e−θ⁢yF(y)=1-e^{-\theta y} for y≥0y\geq 0, and thus 1−F⁢(y)=e−θ⁢y1-F(y)=e^{-\theta y}. Proposition 8.9 gives

    f(1)⁢(y)=n⁢(e−θ⁢y)n−1⁢θ⁢e−θ⁢y=n⁢θ⁢e−n⁢θ⁢y,y≥0,f_{(1)}(y)=n\,\bigl{(}e^{-\theta y}\bigr{)}^{n-1}\,\theta e^{-\theta y}=n% \theta\,e^{-n\theta y},\qquad y\geq 0,

    which is the exponential density with rate n⁢θn\theta. Directly, the minimum exceeds yy if and only if every observation does, and by independence

    ℙθ⁢(X(1)>y)=∏i=1nℙθ⁢(Xi>y)=(e−θ⁢y)n=e−n⁢θ⁢y,\mathbb{P}_{\theta}(X_{(1)}>y)=\prod_{i=1}^{n}\mathbb{P}_{\theta}(X_{i}>y)=% \bigl{(}e^{-\theta y}\bigr{)}^{n}=e^{-n\theta y},

    which is the probability that an exponential random variable with rate n⁢θn\theta exceeds yy. The two computations agree.

  2. 2.
    ​

    From table A.2, an exponential random variable with rate n⁢θn\theta has mean 1/(n⁢θ)1/(n\theta) and variance 1/(n⁢θ)21/(n\theta)^{2}. Thus 𝔼θ⁢(n⁢X(1))=n/(n⁢θ)=1/θ\mathbb{E}_{\theta}(nX_{(1)})=n/(n\theta)=1/\theta, so n⁢X(1)nX_{(1)} is unbiased for 1/θ1/\theta, and Varθ(n⁢X(1))=n2/(n⁢θ)2=1/θ2\mathop{\mathrm{Var}}\nolimits_{\theta}(nX_{(1)})=n^{2}/(n\theta)^{2}=1/\theta% ^{2}.

  3. 3.
    ​

    By lemma 2.2 the sample mean has Varθ(X¯)=1/(n⁢θ2)\mathop{\mathrm{Var}}\nolimits_{\theta}(\bar{X})=1/(n\theta^{2}), and thus we find that Varθ(n⁢X(1))\mathop{\mathrm{Var}}\nolimits_{\theta}(nX_{(1)}) equals n⁢Varθ(X¯)n\mathop{\mathrm{Var}}\nolimits_{\theta}(\bar{X}), so that the estimator based on the minimum has nn times the variance of the sample mean. Worse, its variance does not decrease at all as nn grows. Indeed, by the first part n⁢X(1)nX_{(1)} has the exponential distribution with rate θ\theta for every nn, and thus its sampling distribution does not change with the sample size, and n⁢X(1)nX_{(1)} is not a consistent estimator for 1/θ1/\theta. It is not a function of ∑iXi\sum_{i}X_{i} either: for n=2n=2 the data sets (1,3)(1,3) and (2,2)(2,2) have the same sum but minima 11 and 22. The estimator discards the information which the sufficient statistic retains, and lecture 10 shows how conditioning on a sufficient statistic repairs such an estimator.

∎

Solution to exercise 8.6.
  1. 1.
    ​

    Since X1X_{1} and X2X_{2} are both Bernoulli with success probability θ\theta, we have 𝔼θ⁢(X1−X2)=θ−θ=0\mathbb{E}_{\theta}(X_{1}-X_{2})=\theta-\theta=0 for all θ∈(0,1)\theta\in(0,1) and thus g⁢(T)=X1−X2g(T)=X_{1}-X_{2} is an unbiased estimator for zero. The variable g⁢(T)g(T) equals 11 on the event {X1=1,X2=0}\{X_{1}=1,X_{2}=0\} and, since the two events are independent, this happens with probability θ⁢(1−θ)>0\theta(1-\theta)>0. Thus we have ℙθ⁢(g⁢(T)=0)≤1−θ⁢(1−θ)<1\mathbb{P}_{\theta}\bigl{(}g(T)=0\bigr{)}\leq 1-\theta(1-\theta)<1.

  2. 2.
    ​

    The function gg satisfies 𝔼θ⁢(g⁢(T))=0\mathbb{E}_{\theta}\bigl{(}g(T)\bigr{)}=0 for all θ\theta, but g⁢(T)g(T) does not equal zero with probability one, and thus the condition from definition 8.8 is violated: the whole sample is not complete. Nevertheless, the statistic TT is sufficient, because once the whole sample is given, there is no variability left. This completes the proof.

∎

C.7 Rao–Blackwell and Lehmann–Scheffe

Solution to exercise 10.1.
  1. 1.
    ​

    Since UU takes only the values 0 and 11, we have 𝔼θ⁢(U)=ℙθ⁢(X1=0)=e−θ\mathbb{E}_{\theta}(U)=\mathbb{P}_{\theta}(X_{1}=0)=e^{-\theta} for all θ\theta and thus UU is an unbiased estimator for e−θe^{-\theta}. The indicator function is a Bernoulli random variable with success probability e−θe^{-\theta}, and thus we have Varθ(U)=e−θ⁢(1−e−θ)\mathop{\mathrm{Var}}\nolimits_{\theta}(U)=e^{-\theta}(1-e^{-\theta}).

  2. 2.
    ​

    Let T=X1+RT=X_{1}+R where R=∑i=2nXiR=\sum_{i=2}^{n}X_{i}. By the given fact, RR is Poisson with parameter (n−1)⁢θ(n-1)\theta and TT is Poisson with parameter n⁢θn\theta, and RR is independent of X1X_{1}. For 0≤x≤t0\leq x\leq t, the event {X1=x,T=t}\{X_{1}=x,\ T=t\} is the event {X1=x,R=t−x}\{X_{1}=x,\ R=t-x\} and thus we get

    ℙθ⁢(X1=x|T=t)\displaystyle\mathbb{P}_{\theta}(X_{1}=x\mskip 1.0mu|\mskip 1.0muT=t)
    =ℙθ⁢(X1=x)⁢ℙθ⁢(R=t−x)ℙθ⁢(T=t)\displaystyle=\frac{\mathbb{P}_{\theta}(X_{1}=x)\,\mathbb{P}_{\theta}(R=t-x)}{% \mathbb{P}_{\theta}(T=t)}
    =e−θ⁢θxx!⋅e−(n−1)⁢θ⁢((n−1)⁢θ)t−x(t−x)!⋅t!e−n⁢θ⁢(n⁢θ)t\displaystyle=\frac{e^{-\theta}\theta^{x}}{x!}\cdot\frac{e^{-(n-1)\theta}\bigl% {(}(n-1)\theta\bigr{)}^{t-x}}{(t-x)!}\cdot\frac{t!}{e^{-n\theta}(n\theta)^{t}}
    =t!x!⁢(t−x)!⋅(n−1)t−xnt=(tx)⁢(1n)x⁢(1−1n)t−x.\displaystyle=\frac{t!}{x!\,(t-x)!}\cdot\frac{(n-1)^{t-x}}{n^{t}}=\binom{t}{x}% \Bigl{(}\frac{1}{n}\Bigr{)}^{x}\Bigl{(}1-\frac{1}{n}\Bigr{)}^{t-x}.

    As expected, the parameter θ\theta cancels since TT is sufficient. Given T=tT=t, the first observation can be seen as the number of successes in tt independent trials with success probability 1/n1/n, and in the model each of the tt counted events is equally likely to have been any of the nn observations.

  3. 3.
    ​

    The conditional expectation of an indicator function is a conditional probability. Thus, using the second part of the question, we find

    V=𝔼θ⁢(U|T)=ℙθ⁢(X1=0|T)=(T0)⁢(1−1n)T=(1−1n)T.V=\mathbb{E}_{\theta}(U\mskip 1.0mu|\mskip 1.0muT)=\mathbb{P}_{\theta}(X_{1}=0% \mskip 1.0mu|\mskip 1.0muT)=\binom{T}{0}\Bigl{(}1-\frac{1}{n}\Bigr{)}^{T}=% \Bigl{(}1-\frac{1}{n}\Bigr{)}^{T}.
  4. 4.
    ​

    Using the probability generating function with s=1−1/ns=1-1/n, we find

    𝔼θ⁢(V)=𝔼θ⁢((1−1n)T)=exp⁡(n⁢θ⁢(1−1n−1))=e−θ.\mathbb{E}_{\theta}(V)=\mathbb{E}_{\theta}\Bigl{(}\Bigl{(}1-\frac{1}{n}\Bigr{)% }^{T}\Bigr{)}=\exp\Bigl{(}n\theta\Bigl{(}1-\frac{1}{n}-1\Bigr{)}\Bigr{)}=e^{-% \theta}.

    This agrees with the second statement of the Rao–Blackwell theorem, theorem 10.1.

  5. 5.
    ​

    From example 10.6 we know that the sum TT is a complete and sufficient statistic, and VV is a function of TT which is unbiased for e−θe^{-\theta}. Thus, by the Lehmann–Scheffe theorem, theorem 10.4, VV is the UMVUE for e−θe^{-\theta}. The MLE e−X¯=e−T/ne^{-\bar{X}}=e^{-T/n} is also a function of TT, but in exercise 5.7 we have seen that 𝔼θ⁢(e−X¯)=exp⁡(n⁢θ⁢(e−1/n−1))>e−θ\mathbb{E}_{\theta}(e^{-\bar{X}})=\exp\bigl{(}n\theta(e^{-1/n}-1)\bigr{)}>e^{-\theta}, and thus it is biased and cannot be the UMVUE. Nevertheless, for large nn, the two estimators are close since (1−1/n)T=exp⁡(T⁢log⁡(1−1/n))(1-1/n)^{T}=\exp\bigl{(}T\log(1-1/n)\bigr{)} and log⁡(1−1/n)=−1/n−1/(2⁢n2)−⋯\log(1-1/n)=-1/n-1/(2n^{2})-\dotsb and thus V≈e−X¯⁢e−X¯/(2⁢n)V\approx e^{-\bar{X}}e^{-\bar{X}/(2n)}: the UMVUE is obtained by applying a small downward correction to the MLE in order to compensate for the upward bias.

∎

Solution to exercise 10.2.
  1. 1.
    ​

    The density θ⁢e−θ⁢x\theta e^{-\theta x} is that of the exponential distribution with rate θ\theta, whose mean is 1/θ1/\theta by table A.2, and thus X1X_{1} is unbiased for 1/θ1/\theta. The joint density of the sample is

    f⁢(x1,…,xn;θ)=∏i=1nθ⁢e−θ⁢xi=θn⁢exp⁡(−θ⁢∑i=1nxi)f(x_{1},\dots,x_{n};\theta)=\prod_{i=1}^{n}\theta e^{-\theta x_{i}}=\theta^{n}% \exp\Bigl{(}-\theta\sum_{i=1}^{n}x_{i}\Bigr{)}

    for all xi≥0x_{i}\geq 0, which depends on the data only through ∑ixi\sum_{i}x_{i}, and thus TT is sufficient by the factorisation theorem, theorem 8.2, with h≡1h\equiv 1.

  2. 2.
    ​

    The random variables X1X_{1} and RR are independent, X1X_{1} has density f⁢(x;θ)=θ⁢e−θ⁢xf(x;\theta)=\theta e^{-\theta x} and RR has the quoted gamma density with k=n−1k=n-1, and thus the pair (X1,R)(X_{1},R) has joint density f⁢(x;θ)⁢fn−1⁢(r)f(x;\theta)f_{n-1}(r) for x,r≥0x,r\geq 0. The map (x,r)↦(x,s)=(x,x+r)(x,r)\mapsto(x,s)=(x,x+r) is linear with determinant one, and thus by the change of variable formula, proposition A.19, the pair (X1,T)(X_{1},T) has joint density

    fX1,T⁢(x,s)\displaystyle f_{X_{1},T}(x,s)
    =f⁢(x;θ)⁢fn−1⁢(s−x)=θ⁢e−θ⁢x⋅θn−1⁢(s−x)n−2⁢e−θ⁢(s−x)(n−2)!\displaystyle=f(x;\theta)\,f_{n-1}(s-x)=\theta e^{-\theta x}\cdot\frac{\theta^% {n-1}(s-x)^{n-2}e^{-\theta(s-x)}}{(n-2)!}
    =θn⁢(s−x)n−2⁢e−θ⁢s(n−2)!\displaystyle=\frac{\theta^{n}(s-x)^{n-2}e^{-\theta s}}{(n-2)!}

    for 0<x<s0<x<s, and zero otherwise. Dividing by the density fn⁢(s)=θn⁢sn−1⁢e−θ⁢s/(n−1)!f_{n}(s)=\theta^{n}s^{n-1}e^{-\theta s}/(n-1)! of TT, we obtain the conditional density

    f⁢(x|s)=fX1,T⁢(x,s)fn⁢(s)=(n−1)!(n−2)!⋅(s−x)n−2sn−1=(n−1)⁢(s−x)n−2sn−1f(x\mskip 1.0mu|\mskip 1.0mus)=\frac{f_{X_{1},T}(x,s)}{f_{n}(s)}=\frac{(n-1)!}% {(n-2)!}\cdot\frac{(s-x)^{n-2}}{s^{n-1}}=\frac{(n-1)(s-x)^{n-2}}{s^{n-1}}

    for 0<x<s0<x<s. Both θn\theta^{n} and e−θ⁢se^{-\theta s} have cancelled, and thus the conditional density does not depend on θ\theta, which is the sufficiency of TT seen from definition 8.1. For n=2n=2 the density is 1/s1/s on (0,s)(0,s): given the total of two observations, the first is uniformly distributed over the possible values.

  3. 3.
    ​

    Substituting x=s⁢ux=su, we find

    𝔼θ⁢(X1|T=s)=∫0sx⁢(n−1)⁢(s−x)n−2sn−1⁢dx=(n−1)⁢s⁢∫01u⁢(1−u)n−2⁢du.\mathbb{E}_{\theta}(X_{1}\mskip 1.0mu|\mskip 1.0muT=s)=\int_{0}^{s}x\,\frac{(n% -1)(s-x)^{n-2}}{s^{n-1}}\,\mathrm{d}x=(n-1)\,s\int_{0}^{1}u(1-u)^{n-2}\,% \mathrm{d}u.

    Integrating by parts, with uu differentiated and (1−u)n−2(1-u)^{n-2} integrated, gives

    ∫01u⁢(1−u)n−2⁢du=[−u⁢(1−u)n−1n−1]01+1n−1⁢∫01(1−u)n−1⁢du=1(n−1)⁢n,\int_{0}^{1}u(1-u)^{n-2}\,\mathrm{d}u=\Bigl{[}-\frac{u(1-u)^{n-1}}{n-1}\Bigr{]% }_{0}^{1}+\frac{1}{n-1}\int_{0}^{1}(1-u)^{n-1}\,\mathrm{d}u=\frac{1}{(n-1)n},

    and thus 𝔼θ⁢(X1|T=s)=s/n\mathbb{E}_{\theta}(X_{1}\mskip 1.0mu|\mskip 1.0muT=s)=s/n. Consequently 𝔼θ⁢(X1|T)=T/n=X¯\mathbb{E}_{\theta}(X_{1}\mskip 1.0mu|\mskip 1.0muT)=T/n=\bar{X}: Rao–Blackwellising the single observation X1X_{1}, which has variance 1/θ21/\theta^{2}, gives the sample mean, which has variance 1/(n⁢θ2)1/(n\theta^{2}), and the reduction promised by theorem 10.1 is by a factor of nn.

  4. 4.
    ​

    Since the XiX_{i} are i.i.d. and TT is a symmetric function of them, the pair (Xi,T)(X_{i},T) has the same joint distribution for every ii, and thus the conditional expectations 𝔼θ⁢(Xi|T)\mathbb{E}_{\theta}(X_{i}\mskip 1.0mu|\mskip 1.0muT) are the same random variable for all ii. Summing over ii and using the linearity of conditional expectation, we find

    n⁢𝔼θ⁢(X1|T)=∑i=1n𝔼θ⁢(Xi|T)=𝔼θ⁢(T|T)=T,n\,\mathbb{E}_{\theta}(X_{1}\mskip 1.0mu|\mskip 1.0muT)=\sum_{i=1}^{n}\mathbb{% E}_{\theta}(X_{i}\mskip 1.0mu|\mskip 1.0muT)=\mathbb{E}_{\theta}(T\mskip 1.0mu% |\mskip 1.0muT)=T,

    and thus 𝔼θ⁢(X1|T)=T/n\mathbb{E}_{\theta}(X_{1}\mskip 1.0mu|\mskip 1.0muT)=T/n. This argument uses only the symmetry of an i.i.d. sample and none of the exponential form, and so it applies to every model in which ∑iXi\sum_{i}X_{i} is sufficient.

∎

Solution to exercise 10.3.
  1. 1.
    ​

    In exercise 2.2 we have found 𝔼θ⁢(M)=n⁢θ/(n+1)\mathbb{E}_{\theta}(M)=n\theta/(n+1) and thus

    𝔼θ⁢(θ^)=n+1n⋅n⁢θn+1=θ\mathbb{E}_{\theta}(\hat{\theta})=\frac{n+1}{n}\cdot\frac{n\theta}{n+1}=\theta

    for all θ\theta. Thus, θ^\hat{\theta} is unbiased.

  2. 2.
    ​

    From exercise 8.2 we know that the maximum MM is sufficient for θ\theta and by the second statement of proposition 10.5 it is complete. Since θ^\hat{\theta} is a function of MM and is unbiased for θ\theta, we can use the Lehmann–Scheffe theorem, theorem 10.4, to conclude that this is the UMVUE for θ\theta.

  3. 3.
    ​

    From the solution to exercise 2.2 we know Varθ(M)=n⁢θ2/((n+2)⁢(n+1)2)\mathop{\mathrm{Var}}\nolimits_{\theta}(M)=n\theta^{2}/\bigl{(}(n+2)(n+1)^{2}% \bigr{)} and thus

    Varθ(θ^)=(n+1n)2⋅n⁢θ2(n+2)⁢(n+1)2=θ2n⁢(n+2).\mathop{\mathrm{Var}}\nolimits_{\theta}(\hat{\theta})=\Bigl{(}\frac{n+1}{n}% \Bigr{)}^{2}\cdot\frac{n\theta^{2}}{(n+2)(n+1)^{2}}=\frac{\theta^{2}}{n(n+2)}.

    The variance of a single observation is θ2/12\theta^{2}/12, as shown in table A.2, and thus, using lemma 2.2, we find

    Varθ(θ~)=4⁢Varθ(X¯)=4⁢θ212⁢n=θ23⁢n.\mathop{\mathrm{Var}}\nolimits_{\theta}(\tilde{\theta})=4\,\mathop{\mathrm{Var% }}\nolimits_{\theta}(\bar{X})=\frac{4\theta^{2}}{12n}=\frac{\theta^{2}}{3n}.

    The ratio Varθ(θ^)/Varθ(θ~)=3/(n+2)\mathop{\mathrm{Var}}\nolimits_{\theta}(\hat{\theta})/\mathop{\mathrm{Var}}% \nolimits_{\theta}(\tilde{\theta})=3/(n+2), which is always less than or equal to one for all n≥1n\geq 1, equals one only for n=1n=1, where both estimators equal 2⁢X12X_{1}. For large nn, the variance of the UMVUE is of order 1/n21/n^{2}, compared to the order 1/n1/n for 2⁢X¯2\bar{X}, similar to the result in exercise 2.2.

  4. 4.
    ​

    From theorem 10.1, the Rao–Blackwell theorem, we know that the conditional expectation 𝔼θ⁢(θ~|M)\mathbb{E}_{\theta}(\tilde{\theta}\mskip 1.0mu|\mskip 1.0muM) is a random variable, which is a function of MM and is unbiased for θ\theta. Since MM is complete, the first step of the proof of theorem 10.4 shows that there can only be one such function and thus we can conclude 𝔼θ⁢(θ~|M)=(n+1)⁢M/n\mathbb{E}_{\theta}(\tilde{\theta}\mskip 1.0mu|\mskip 1.0muM)=(n+1)M/n with probability one. For n=1n=1 we have M=X1M=X_{1} and θ~=2⁢X1=2⁢M\tilde{\theta}=2X_{1}=2M, which coincides with the formula (n+1)⁢M/n=2⁢M(n+1)M/n=2M.

∎

Solution to exercise 10.4.
  1. 1.
    ​

    Using the definition of the exponential family, we can write the density as

    f⁢(x;μ,σ2)=12⁢π⁢σ2⁢exp⁡(−(x−μ)22⁢σ2)=e−μ2/(2⁢σ2)2⁢π⁢σ2⁢exp⁡(μσ2⁢x−12⁢σ2⁢x2),f(x;\mu,\sigma^{2})=\frac{1}{\sqrt{2\pi\sigma^{2}}}\exp\Bigl{(}-\frac{(x-\mu)^% {2}}{2\sigma^{2}}\Bigr{)}=\frac{e^{-\mu^{2}/(2\sigma^{2})}}{\sqrt{2\pi\sigma^{% 2}}}\exp\Bigl{(}\frac{\mu}{\sigma^{2}}\,x-\frac{1}{2\sigma^{2}}\,x^{2}\Bigr{)},

    which is of the form given in definition 7.8 with h⁢(x)=1h(x)=1, natural statistic T⁢(x)=(x,x2)T(x)=(x,x^{2}) and natural parameter

    η⁢(μ,σ2)=(μσ2,−12⁢σ2).\eta(\mu,\sigma^{2})=\Bigl{(}\frac{\mu}{\sigma^{2}},\ -\frac{1}{2\sigma^{2}}% \Bigr{)}.
  2. 2.
    ​

    As (μ,σ2)(\mu,\sigma^{2}) varies over ℝ×(0,∞)\mathbb{R}\times(0,\infty), the second component η2=−1/(2⁢σ2)\eta_{2}=-1/(2\sigma^{2}) can take any value in (−∞,0)(-\infty,0). For fixed σ2\sigma^{2}, the first component η1=μ/σ2\eta_{1}=\mu/\sigma^{2} can take any value in ℝ\mathbb{R} as μ\mu varies. Thus, the range of η\eta is the open half-plane ℝ×(−∞,0)\mathbb{R}\times(-\infty,0), which is an open, non-empty subset of ℝ2\mathbb{R}^{2} and thus the family has full rank in the sense of definition 7.10.

  3. 3.
    ​

    From corollary 8.3 and the first statement of proposition 10.5 we know that the natural statistic (∑iXi,∑iXi2)\bigl{(}\sum_{i}X_{i},\sum_{i}X_{i}^{2}\bigr{)} of the sample is complete and sufficient. From the pair (∑iXi,∑iXi2)\bigl{(}\sum_{i}X_{i},\sum_{i}X_{i}^{2}\bigr{)} we can compute X¯=∑iXi/n\bar{X}=\sum_{i}X_{i}/n and S2=(∑iXi2−n⁢X¯2)/(n−1)S^{2}=\bigl{(}\sum_{i}X_{i}^{2}-n\bar{X}^{2}\bigr{)}/(n-1), and conversely ∑iXi=n⁢X¯\sum_{i}X_{i}=n\bar{X} and ∑iXi2=(n−1)⁢S2+n⁢X¯2\sum_{i}X_{i}^{2}=(n-1)S^{2}+n\bar{X}^{2}, and thus the map between the two pairs is one-to-one. Since a one-to-one function of a complete sufficient statistic is itself complete and sufficient, the pair (X¯,S2)(\bar{X},S^{2}) is complete and sufficient.

∎

Solution to exercise 10.5.
  1. 1.
    ​

    Let Zi=Xi−μZ_{i}=X_{i}-\mu, so that Z1,…,ZnZ_{1},\dots,Z_{n} are i.i.d. N⁢(0,σ2)N(0,\sigma^{2}) whatever the value of μ\mu. Since subtracting the same constant from every observation shifts the maximum and the minimum by that constant, we have

    R=maxi⁡Xi−mini⁡Xi=maxi⁡Zi−mini⁡Zi,R=\max_{i}X_{i}-\min_{i}X_{i}=\max_{i}Z_{i}-\min_{i}Z_{i},

    and the right-hand side is a function of random variables whose joint distribution does not involve μ\mu. Thus the distribution of RR is the same under every ℙμ\mathbb{P}_{\mu}, and RR is ancillary for μ\mu in the sense of definition 10.8.

  2. 2.
    ​

    In the proof of corollary 10.10 we have seen that, for known σ2\sigma^{2}, the normal distribution is a full-rank exponential family with natural statistic T⁢(x)=xT(x)=x, and thus X¯\bar{X} is complete and sufficient for μ\mu by corollary 8.3 and proposition 10.5. Since RR is ancillary by the first part, Basu’s theorem, theorem 10.9, shows that X¯\bar{X} and RR are independent.

  3. 3.
    ​

    Now let μ\mu be known and σ2\sigma^{2} unknown, and write Xi=μ+σ⁢WiX_{i}=\mu+\sigma W_{i} with W1,…,WnW_{1},\dots,W_{n} i.i.d. standard normal. Then R=σ⁢(maxi⁡Wi−mini⁡Wi)R=\sigma\,(\max_{i}W_{i}-\min_{i}W_{i}), and thus for example 𝔼σ2⁢(R)=σ⁢𝔼⁢(maxi⁡Wi−mini⁡Wi)\mathbb{E}_{\sigma^{2}}(R)=\sigma\,\mathbb{E}(\max_{i}W_{i}-\min_{i}W_{i}) is proportional to σ\sigma, and this expectation is strictly positive rather than zero, since W1W_{1} and W2W_{2} are independent and continuous and thus differ with probability one. The distribution of RR changes with the parameter, and RR is not ancillary for σ2\sigma^{2}. The scaled range R/σR/\sigma does have a distribution free of σ2\sigma^{2}, but it is not a statistic in the sense of definition 1.3, because computing it requires the unknown σ\sigma. Quantities of this kind, whose distribution is known although they involve the parameter, return in lecture 25, where they are used to build confidence intervals.

∎

C.8 Fisher Information

Solution to exercise 11.1.
  1. 1.
    ​

    The score for the observation is given by ℓ1′⁢(θ)=X1/θ−(1−X1)/(1−θ)\ell_{1}^{\prime}(\theta)=X_{1}/\theta-(1-X_{1})/(1-\theta). This takes values 1/θ1/\theta if X1=1X_{1}=1 (with probability θ\theta) and −1/(1−θ)-1/(1-\theta) if X1=0X_{1}=0 (with probability 1−θ1-\theta). Thus we have

    𝔼θ⁢(ℓ1′⁢(θ))=θ⋅1θ+(1−θ)⋅(−11−θ)=1−1=0,\mathbb{E}_{\theta}\bigl{(}\ell_{1}^{\prime}(\theta)\bigr{)}=\theta\cdot\frac{% 1}{\theta}+(1-\theta)\cdot\Bigl{(}-\frac{1}{1-\theta}\Bigr{)}=1-1=0,

    as given in lemma 11.1.

  2. 2.
    ​

    Combining the two fractions we get

    ℓ1′⁢(θ)=X1⁢(1−θ)−(1−X1)⁢θθ⁢(1−θ)=X1−θθ⁢(1−θ),\ell_{1}^{\prime}(\theta)=\frac{X_{1}(1-\theta)-(1-X_{1})\theta}{\theta(1-% \theta)}=\frac{X_{1}-\theta}{\theta(1-\theta)},

    which is a constant multiple of X1−θX_{1}-\theta. Since Varθ(X1)=θ⁢(1−θ)\mathop{\mathrm{Var}}\nolimits_{\theta}(X_{1})=\theta(1-\theta) we find

    Varθ(ℓ1′⁢(θ))=θ⁢(1−θ)θ2⁢(1−θ)2=1θ⁢(1−θ).\mathop{\mathrm{Var}}\nolimits_{\theta}\bigl{(}\ell_{1}^{\prime}(\theta)\bigr{% )}=\frac{\theta(1-\theta)}{\theta^{2}(1-\theta)^{2}}=\frac{1}{\theta(1-\theta)}.

    This is the value of ℐ︀⁢(θ)\mathcal{I}(\theta) in example 11.5 and thus, by theorem 11.3, the correct value.

  3. 3.
    ​

    From proposition 11.4 we know ℐ︀n⁢(θ)=n/(θ⁢(1−θ))\mathcal{I}_{n}(\theta)=n/\bigl{(}\theta(1-\theta)\bigr{)} and thus 1/ℐ︀n⁢(θ)=θ⁢(1−θ)/n1/\mathcal{I}_{n}(\theta)=\theta(1-\theta)/n. By lemma 2.2 the sample mean has Varθ(X¯)=Varθ(X1)/n=θ⁢(1−θ)/n\mathop{\mathrm{Var}}\nolimits_{\theta}(\bar{X})=\mathop{\mathrm{Var}}% \nolimits_{\theta}(X_{1})/n=\theta(1-\theta)/n and thus the variance of X¯\bar{X} equals the reciprocal of the information, as it did for the normal mean in example 11.7.

∎

Solution to exercise 11.2.
  1. 1.
    ​

    The log-likelihood of a single observation is ℓ1⁢(θ)=−2⁢log⁡θ+log⁡X1−X1/θ\ell_{1}(\theta)=-2\log\theta+\log X_{1}-X_{1}/\theta and thus the score is given by

    ℓ1′⁢(θ)=−2θ+X1θ2.\ell_{1}^{\prime}(\theta)=-\frac{2}{\theta}+\frac{X_{1}}{\theta^{2}}.

    From table A.2 or from the solution to exercise 5.2 we know 𝔼θ⁢(X1)=2⁢θ\mathbb{E}_{\theta}(X_{1})=2\theta and thus 𝔼θ⁢(ℓ1′⁢(θ))=−2/θ+2⁢θ/θ2=0\mathbb{E}_{\theta}\bigl{(}\ell_{1}^{\prime}(\theta)\bigr{)}=-2/\theta+2\theta% /\theta^{2}=0.

  2. 2.
    ​

    Taking derivatives again, we find ℓ1′′⁢(θ)=2/θ2−2⁢X1/θ3\ell_{1}^{\prime\prime}(\theta)=2/\theta^{2}-2X_{1}/\theta^{3} and thus

    ℐ︀⁢(θ)=−𝔼θ⁢(ℓ1′′⁢(θ))=−2θ2+2⋅2⁢θθ3=2θ2.\mathcal{I}(\theta)=-\mathbb{E}_{\theta}\bigl{(}\ell_{1}^{\prime\prime}(\theta% )\bigr{)}=-\frac{2}{\theta^{2}}+\frac{2\cdot 2\theta}{\theta^{3}}=\frac{2}{% \theta^{2}}.

    On the other hand, since the score is X1/θ2X_{1}/\theta^{2} plus a constant, we can find the variance as

    Varθ(ℓ1′⁢(θ))=Varθ(X1)θ4=2⁢θ2θ4=2θ2.\mathop{\mathrm{Var}}\nolimits_{\theta}\bigl{(}\ell_{1}^{\prime}(\theta)\bigr{% )}=\frac{\mathop{\mathrm{Var}}\nolimits_{\theta}(X_{1})}{\theta^{4}}=\frac{2% \theta^{2}}{\theta^{4}}=\frac{2}{\theta^{2}}.

    Both approaches give the same result.

  3. 3.
    ​

    Since the density can be written as

    f⁢(x;θ)=x⁢exp⁡(−xθ−2⁢log⁡θ),f(x;\theta)=x\,\exp\Bigl{(}-\frac{x}{\theta}-2\log\theta\Bigr{)},

    we have h⁢(x)=xh(x)=x, T⁢(x)=xT(x)=x, η⁢(θ)=−1/θ\eta(\theta)=-1/\theta and B⁢(θ)=2⁢log⁡θB(\theta)=2\log\theta. Thus, the model is an exponential family as in definition 7.1. The natural parameter takes values η<0\eta<0 and the cumulant function is

    K⁢(η)=log⁢∫0∞x⁢eη⁢x⁢dx=log⁡1η2=−2⁢log⁡(−η),K(\eta)=\log\int_{0}^{\infty}x\,e^{\eta x}\,\mathrm{d}x=\log\frac{1}{\eta^{2}}% =-2\log(-\eta),

    where K′⁢(η)=−2/η=2⁢θK^{\prime}(\eta)=-2/\eta=2\theta is the mean and K′′⁢(η)=2/η2=2⁢θ2K^{\prime\prime}(\eta)=2/\eta^{2}=2\theta^{2} is the variance, as predicted by theorem 7.7. Since η′⁢(θ)=1/θ2\eta^{\prime}(\theta)=1/\theta^{2}, we find by proposition 11.9 ℐ︀⁢(θ)=η′⁢(θ)2⁢K′′⁢(η⁢(θ))=2⁢θ2/θ4=2/θ2\mathcal{I}(\theta)=\eta^{\prime}(\theta)^{2}K^{\prime\prime}\bigl{(}\eta(% \theta)\bigr{)}=2\theta^{2}/\theta^{4}=2/\theta^{2}. This is the same value as we obtained above.

  4. 4.
    ​

    From proposition 11.4 we know that the information of the sample is ℐ︀n⁢(θ)=2⁢n/θ2\mathcal{I}_{n}(\theta)=2n/\theta^{2} and thus 1/ℐ︀n⁢(θ)=θ2/(2⁢n)=MSE(θ^)1/\mathcal{I}_{n}(\theta)=\theta^{2}/(2n)=\mathop{\mathrm{MSE}}\nolimits(\hat{% \theta}). Since θ^\hat{\theta} is unbiased, its mean squared error equals its variance and thus the variance of the MLE is the reciprocal of the Fisher information for all sample sizes, as it was for the sample mean in exercise 11.1 and example 11.7. From lecture 13 we know that this is the smallest variance an unbiased estimator can have.

∎

Solution to exercise 11.3.
  1. 1.
    ​

    The log-likelihood of a single observation is ℓ1⁢(p)=(X1−1)⁢log⁡(1−p)+log⁡p\ell_{1}(p)=(X_{1}-1)\log(1-p)+\log p and thus the score is given by

    ℓ1′⁢(p)=1p−X1−11−p.\ell_{1}^{\prime}(p)=\frac{1}{p}-\frac{X_{1}-1}{1-p}.

    Using 𝔼p⁢(X1−1)=1/p−1=(1−p)/p\mathbb{E}_{p}(X_{1}-1)=1/p-1=(1-p)/p we find

    𝔼p⁢(ℓ1′⁢(p))=1p−(1−p)/p1−p=1p−1p=0.\mathbb{E}_{p}\bigl{(}\ell_{1}^{\prime}(p)\bigr{)}=\frac{1}{p}-\frac{(1-p)/p}{% 1-p}=\frac{1}{p}-\frac{1}{p}=0.
  2. 2.
    ​

    Taking derivatives again, we find ℓ1′′⁢(p)=−1/p2−(X1−1)/(1−p)2\ell_{1}^{\prime\prime}(p)=-1/p^{2}-(X_{1}-1)/(1-p)^{2} and taking expectations we get

    ℐ︀⁢(p)\displaystyle\mathcal{I}(p)
    =−𝔼p⁢(ℓ1′′⁢(p))=1p2+(1−p)/p(1−p)2\displaystyle=-\mathbb{E}_{p}\bigl{(}\ell_{1}^{\prime\prime}(p)\bigr{)}=\frac{% 1}{p^{2}}+\frac{(1-p)/p}{(1-p)^{2}}
    =1p2+1p⁢(1−p)=(1−p)+pp2⁢(1−p)=1p2⁢(1−p).\displaystyle=\frac{1}{p^{2}}+\frac{1}{p(1-p)}=\frac{(1-p)+p}{p^{2}(1-p)}=% \frac{1}{p^{2}(1-p)}.

    From proposition 11.4 we know that the sample has information ℐ︀n⁢(p)=n/(p2⁢(1−p))\mathcal{I}_{n}(p)=n/\bigl{(}p^{2}(1-p)\bigr{)}.

  3. 3.
    ​

    Since the score is −X1/(1−p)-X_{1}/(1-p) plus a constant, we find

    Varp(ℓ1′⁢(p))=Varp(X1)(1−p)2=(1−p)/p2(1−p)2=1p2⁢(1−p).\mathop{\mathrm{Var}}\nolimits_{p}\bigl{(}\ell_{1}^{\prime}(p)\bigr{)}=\frac{% \mathop{\mathrm{Var}}\nolimits_{p}(X_{1})}{(1-p)^{2}}=\frac{(1-p)/p^{2}}{(1-p)% ^{2}}=\frac{1}{p^{2}(1-p)}.

    This agrees with the result from the previous part.

∎

Solution to exercise 11.4.
  1. 1.
    ​

    Since K⁢(η)=eηK(\eta)=e^{\eta} we have K′′⁢(η)=eηK^{\prime\prime}(\eta)=e^{\eta} and using proposition 11.9 we find ℐ︀η⁢(η)=eη=θ\mathcal{I}_{\eta}(\eta)=e^{\eta}=\theta. This is the variance of the natural statistic T⁢(X)=XT(X)=X, as expected.

  2. 2.
    ​

    Using proposition 11.10 with η\eta (the original parameter) and θ=g⁢(η)=eη\theta=g(\eta)=e^{\eta} (the new parameter), we find g′⁢(η)=eη=θg^{\prime}(\eta)=e^{\eta}=\theta. Thus we have

    ℐ︀θ⁢(θ)=ℐ︀η⁢(η)g′⁢(η)2=θθ2=1θ,\mathcal{I}_{\theta}(\theta)=\frac{\mathcal{I}_{\eta}(\eta)}{g^{\prime}(\eta)^% {2}}=\frac{\theta}{\theta^{2}}=\frac{1}{\theta},

    the value from example 11.6.

  3. 3.
    ​

    We have p0=g⁢(θ)=e−θp_{0}=g(\theta)=e^{-\theta} with g′⁢(θ)=−e−θg^{\prime}(\theta)=-e^{-\theta} and using proposition 11.10 we find

    ℐ︀p0⁢(p0)=1/θe−2⁢θ=e2⁢θθ,\mathcal{I}_{p_{0}}(p_{0})=\frac{1/\theta}{e^{-2\theta}}=\frac{e^{2\theta}}{% \theta},

    or using the notation p0p_{0} we find ℐ︀p0⁢(p0)=1/(p02⁢log⁡(1/p0))\mathcal{I}_{p_{0}}(p_{0})=1/\bigl{(}p_{0}^{2}\log(1/p_{0})\bigr{)}. As θ→∞\theta\to\infty, the information goes to infinity. This makes sense, since p0=e−θp_{0}=e^{-\theta} converges to zero, and the information measures accuracy on the absolute scale of the parameter: a probability which is close to zero is estimated with small absolute error by any reasonable estimator, and the approximate variance 1/ℐ︀n⁢(p0)=θ⁢e−2⁢θ/n1/\mathcal{I}_{n}(p_{0})=\theta e^{-2\theta}/n from lecture 17 is tiny, even if the error relative to the size of p0p_{0} is not small.

∎

Solution to exercise 11.5.
  1. 1.
    ​

    The map g⁢(θ)=1/θg(\theta)=1/\theta maps (0,∞)(0,\infty) to (0,∞)(0,\infty) and g′⁢(θ)=−1/θ2≠0g^{\prime}(\theta)=-1/\theta^{2}\neq 0. From example 11.8 we know that ℐ︀θ⁢(θ)=1/θ2\mathcal{I}_{\theta}(\theta)=1/\theta^{2}. Thus, by proposition 11.10, we have

    ℐ︀μ⁢(μ)=ℐ︀θ⁢(θ)g′⁢(θ)2=1/θ21/θ4=θ2=1μ2.\mathcal{I}_{\mu}(\mu)=\frac{\mathcal{I}_{\theta}(\theta)}{g^{\prime}(\theta)^% {2}}=\frac{1/\theta^{2}}{1/\theta^{4}}=\theta^{2}=\frac{1}{\mu^{2}}.

    This is the value given in section 11.6.

  2. 2.
    ​

    Using proposition 11.4, in the μ\mu-parametrisation, the information of the sample about μ\mu is given by n⁢ℐ︀μ⁢(μ)=n/μ2n\mathcal{I}_{\mu}(\mu)=n/\mu^{2}. From table A.2 we know that Varμ(X1)=1/θ2=μ2\mathop{\mathrm{Var}}\nolimits_{\mu}(X_{1})=1/\theta^{2}=\mu^{2} and thus we have Varμ(X¯)=μ2/n\mathop{\mathrm{Var}}\nolimits_{\mu}(\bar{X})=\mu^{2}/n, by lemma 2.2. The variance of the unbiased estimator X¯\bar{X} for μ\mu equals the reciprocal μ2/n\mu^{2}/n of the information of the sample, and from lecture 13 we know that no unbiased estimator can be better.

∎

Solution to exercise 11.6.

The right-hand side of (11.1) with k=1k=1 is an integral over the whole real line, but f⁢(x;θ)=0f(x;\theta)=0 and thus ∂f/∂θ=0\partial f/\partial\theta=0 for x<0x<0 and for x>θx>\theta; since x=θx=\theta is a single point, it does not contribute to the integral; only 0≤x≤θ0\leq x\leq\theta, where f⁢(x;θ)=1/θf(x;\theta)=1/\theta, is left. We have

∫0θ∂∂θ⁢f⁢(x;θ)⁢dx=∫0θ(−1θ2)⁢dx=−1θ,\int_{0}^{\theta}\frac{\partial}{\partial\theta}f(x;\theta)\,\mathrm{d}x=\int_% {0}^{\theta}\Bigl{(}-\frac{1}{\theta^{2}}\Bigr{)}\mathrm{d}x=-\frac{1}{\theta},

while the left-hand side can be found as

dd⁢θ⁢∫0θf⁢(x;θ)⁢dx=dd⁢θ⁢ 1=0.\frac{\mathrm{d}}{\mathrm{d}\theta}\int_{0}^{\theta}f(x;\theta)\,\mathrm{d}x=% \frac{\mathrm{d}}{\mathrm{d}\theta}\,1=0.

The difference between the two sides is 0−(−1/θ)=1/θ0-(-1/\theta)=1/\theta, which is the value f⁢(θ;θ)=1/θf(\theta;\theta)=1/\theta of the density at the upper limit of the integral. The derivative of the integral picks up this boundary term (from the moving upper limit) in addition to the derivative of the integrand, but the differentiation under the integral sign ignores this boundary term. ∎

C.9 The Cramer–Rao Inequality

Solution to exercise 13.1.
  1. 1.
    ​

    For y≥0y\geq 0 we have

    ℙθ⁢(Y1≤y)=ℙθ⁢(X1≤y2⁢θ)=1−e−y/2,\mathbb{P}_{\theta}(Y_{1}\leq y)=\mathbb{P}_{\theta}\Bigl{(}X_{1}\leq\frac{y}{% 2\theta}\Bigr{)}=1-e^{-y/2},

    and taking derivatives we find the density 12⁢e−y/2\tfrac{1}{2}e^{-y/2} for y≥0y\geq 0. This is the exponential density with rate 1/21/2, i.e. the Gamma⁢(1,1/2)\text{Gamma}(1,1/2) density from table A.2 and Gamma⁢(1,1/2)\text{Gamma}(1,1/2) is χ22\chi^{2}_{2} by the statement quoted in the question. Thus we have Y1∼χ22Y_{1}\sim\chi^{2}_{2}.

  2. 2.
    ​

    From the first part and definition A.2 we know that each Yi=2⁢θ⁢XiY_{i}=2\theta X_{i} has the distribution of Z2⁢i−12+Z2⁢i2Z_{2i-1}^{2}+Z_{2i}^{2} for independent, standard normally distributed Z2⁢i−1Z_{2i-1} and Z2⁢iZ_{2i}. Since the XiX_{i} are independent, we can assume that all 2⁢n2n are independent normal variables. Thus, 2⁢θ⁢T=∑i=1nYi2\theta T=\sum_{i=1}^{n}Y_{i} has the distribution of Z12+⋯+Z2⁢n2Z_{1}^{2}+\dots+Z_{2n}^{2}, i.e. χ2⁢n2\chi^{2}_{2n} by definition A.2.

  3. 3.
    ​

    The density of T∼Gamma⁢(n,θ)T\sim\text{Gamma}(n,\theta) is given by fT⁢(t)=θn⁢tn−1⁢e−θ⁢t/Γ⁢(n)f_{T}(t)=\theta^{n}t^{n-1}e^{-\theta t}/\Gamma(n) for t≥0t\geq 0. For V=2⁢θ⁢TV=2\theta T, using proposition A.19 with inverse map t=v/(2⁢θ)t=v/(2\theta) and derivative 1/(2⁢θ)1/(2\theta), we find

    fV⁢(v)=fT⁢(v2⁢θ)⁢12⁢θ=θn⁢(v/(2⁢θ))n−1⁢e−v/2Γ⁢(n)⋅12⁢θ=(1/2)n⁢vn−1⁢e−v/2Γ⁢(n)f_{V}(v)=f_{T}\Bigl{(}\frac{v}{2\theta}\Bigr{)}\,\frac{1}{2\theta}=\frac{% \theta^{n}\bigl{(}v/(2\theta)\bigr{)}^{n-1}e^{-v/2}}{\Gamma(n)}\cdot\frac{1}{2% \theta}=\frac{(1/2)^{n}\,v^{n-1}e^{-v/2}}{\Gamma(n)}

    for v≥0v\geq 0. This is the Gamma⁢(n,1/2)\text{Gamma}(n,1/2) density again, by table A.2. Since n=(2⁢n)/2n=(2n)/2, this is the χ2⁢n2\chi^{2}_{2n} density and it agrees with the result from the previous part.

  4. 4.
    ​

    From definition A.2 we know 𝔼θ⁢(2⁢θ⁢T)=2⁢n\mathbb{E}_{\theta}(2\theta T)=2n and Varθ(2⁢θ⁢T)=4⁢n\mathop{\mathrm{Var}}\nolimits_{\theta}(2\theta T)=4n, and thus 𝔼θ⁢(T)=n/θ\mathbb{E}_{\theta}(T)=n/\theta and Varθ(T)=4⁢n/(4⁢θ2)=n/θ2\mathop{\mathrm{Var}}\nolimits_{\theta}(T)=4n/(4\theta^{2})=n/\theta^{2}. For a single observation, the mean is 1/θ1/\theta and the variance is 1/θ21/\theta^{2}, as found in table A.2, and thus the sum of nn independent copies has mean n/θn/\theta and variance n/θ2n/\theta^{2}, as shown above.

∎

Solution to exercise 13.2.
  1. 1.
    ​

    From example 11.5 we know that ℐ︀n⁢(θ)=n/(θ⁢(1−θ))\mathcal{I}_{n}(\theta)=n/\bigl{(}\theta(1-\theta)\bigr{)}. Thus, the bound from theorem 13.2 is 1/ℐ︀n⁢(θ)=θ⁢(1−θ)/n1/\mathcal{I}_{n}(\theta)=\theta(1-\theta)/n.

  2. 2.
    ​

    The sample mean is an unbiased estimator for θ\theta and using lemma 2.2 we find Varθ(X¯)=θ⁢(1−θ)/n\mathop{\mathrm{Var}}\nolimits_{\theta}(\bar{X})=\theta(1-\theta)/n, since the variance of a Bernoulli distribution is θ⁢(1−θ)\theta(1-\theta). This bound coincides with the bound from the definition, and thus X¯\bar{X} is efficient by definition 13.4, since the exchange identity (13.1) holds because the Bernoulli distribution is a regular exponential family and X¯\bar{X} has finite variance, as noted after lemma 13.1.

  3. 3.
    ​

    The score of the sample in example 11.5 is given by

    ℓ′⁢(θ)=Tnθ−n−Tn1−θ=Tn−n⁢θθ⁢(1−θ)=nθ⁢(1−θ)⁢(X¯−θ),\ell^{\prime}(\theta)=\frac{T_{n}}{\theta}-\frac{n-T_{n}}{1-\theta}=\frac{T_{n% }-n\theta}{\theta(1-\theta)}=\frac{n}{\theta(1-\theta)}\,(\bar{X}-\theta),

    and thus we have X¯−θ=c⁢(θ)⁢ℓ′⁢(θ)\bar{X}-\theta=c(\theta)\,\ell^{\prime}(\theta) with c⁢(θ)=θ⁢(1−θ)/nc(\theta)=\theta(1-\theta)/n. Since g⁢(θ)=θg(\theta)=\theta has g′⁢(θ)=1g^{\prime}(\theta)=1, this is g′⁢(θ)/ℐ︀n⁢(θ)g^{\prime}(\theta)/\mathcal{I}_{n}(\theta) as predicted by proposition 13.5.

  4. 4.
    ​

    For g⁢(θ)=θ/(1−θ)g(\theta)=\theta/(1-\theta) we have g′⁢(θ)=1/(1−θ)2g^{\prime}(\theta)=1/(1-\theta)^{2} and thus using theorem 13.3 we get the bound

    g′⁢(θ)2ℐ︀n⁢(θ)=1(1−θ)4⋅θ⁢(1−θ)n=θn⁢(1−θ)3.\frac{g^{\prime}(\theta)^{2}}{\mathcal{I}_{n}(\theta)}=\frac{1}{(1-\theta)^{4}% }\cdot\frac{\theta(1-\theta)}{n}=\frac{\theta}{n(1-\theta)^{3}}.

    The Bernoulli distribution is an exponential family with natural statistic T⁢(x)=xT(x)=x and μ⁢(θ)=θ\mu(\theta)=\theta. From the discussion after proposition 13.5 we know that only affine functions of the form a+b⁢θa+b\theta can be used to construct an efficient estimator. Since the odds are not an affine function of θ\theta, there can be no efficient estimator for the odds.

  5. 5.
    ​

    Any estimator W=W⁢(X1,…,Xn)W=W(X_{1},\dots,X_{n}) can take at most 2n2^{n} different values, one for each possible value of the sample x∈{0,1}nx\in\{0,1\}^{n}, and the expectation of the estimator is

    𝔼θ⁢(W)=∑x∈{0,1}nW⁢(x)⁢θ∑ixi⁢(1−θ)n−∑ixi,\mathbb{E}_{\theta}(W)=\sum_{x\in\{0,1\}^{n}}W(x)\,\theta^{\sum_{i}x_{i}}(1-% \theta)^{n-\sum_{i}x_{i}},

    which is a polynomial in θ\theta of degree at most nn. A polynomial function is bounded on (0,1)(0,1), whereas θ/(1−θ)→∞\theta/(1-\theta)\to\infty as θ→1\theta\to 1 and thus no choice of WW makes the estimator 𝔼θ⁢(W)\mathbb{E}_{\theta}(W) equal to the odds for all θ\theta. Thus there is no unbiased estimator for the odds, efficient or otherwise.

∎

Solution to exercise 13.3.
  1. 1.
    ​

    From example 11.6 we know that ℐ︀n⁢(θ)=n/θ\mathcal{I}_{n}(\theta)=n/\theta. Thus, the Cramer–Rao bound for unbiased estimators for θ\theta is θ/n\theta/n. Since the Poisson distribution has mean and variance θ\theta, lemma 2.2 shows that 𝔼θ⁢(X¯)=θ\mathbb{E}_{\theta}(\bar{X})=\theta and Varθ(X¯)=θ/n\mathop{\mathrm{Var}}\nolimits_{\theta}(\bar{X})=\theta/n. Thus, X¯\bar{X} is an unbiased estimator and satisfies the bound: it is an efficient estimator for θ\theta, since the exchange identity (13.1) holds because the Poisson distribution is a regular exponential family and X¯\bar{X} has finite variance, as noted after lemma 13.1.

  2. 2.
    ​

    The log-likelihood function is ℓ⁢(θ)=Tn⁢log⁡θ−n⁢θ−∑ilog⁡(Xi!)\ell(\theta)=T_{n}\log\theta-n\theta-\sum_{i}\log(X_{i}!) and thus the score function is

    ℓ′⁢(θ)=Tnθ−n=nθ⁢(X¯−θ).\ell^{\prime}(\theta)=\frac{T_{n}}{\theta}-n=\frac{n}{\theta}\,(\bar{X}-\theta).

    Re-arranging this we find X¯−θ=(θ/n)⁢ℓ′⁢(θ)\bar{X}-\theta=(\theta/n)\,\ell^{\prime}(\theta), which is condition (13.2) with W=X¯W=\bar{X}, g⁢(θ)=θg(\theta)=\theta and c⁢(θ)=θ/nc(\theta)=\theta/n; indeed g′⁢(θ)/ℐ︀n⁢(θ)=1/(n/θ)=θ/ng^{\prime}(\theta)/\mathcal{I}_{n}(\theta)=1/(n/\theta)=\theta/n.

  3. 3.
    ​

    From the discussion after proposition 13.5 we know that in a one-parameter exponential family, only affine functions of the mean μ⁢(θ)\mu(\theta) of the natural statistic allow for efficient estimators. In this case we have μ⁢(θ)=θ\mu(\theta)=\theta and g⁢(θ)=e−θg(\theta)=e^{-\theta} is not an affine function of θ\theta. Thus, no efficient estimator for e−θe^{-\theta} exists. There is no contradiction in this result: example 10.6 shows that (1−1/n)Tn(1-1/n)^{T_{n}} has the smallest variance amongst all unbiased estimators for e−θe^{-\theta}, but this smallest variance is strictly larger than the Cramer–Rao bound. Efficiency is a stronger property than being UMVUE, and in this case the bound cannot be achieved.

∎

Solution to exercise 13.4.
  1. 1.
    ​

    Example 11.8 gives ℐ︀⁢(θ)=1/θ2\mathcal{I}(\theta)=1/\theta^{2}, and thus ℐ︀n⁢(θ)=n/θ2\mathcal{I}_{n}(\theta)=n/\theta^{2} and the bound is 1/ℐ︀n⁢(θ)=θ2/n1/\mathcal{I}_{n}(\theta)=\theta^{2}/n.

  2. 2.
    ​

    Since θ~\tilde{\theta} is unbiased, its mean squared error is its variance, Varθ(θ~)=θ2/(n−2)\mathop{\mathrm{Var}}\nolimits_{\theta}(\tilde{\theta})=\theta^{2}/(n-2). The efficiency is

    effθ⁡(θ~)=θ2/nθ2/(n−2)=n−2n,\operatorname{eff}_{\theta}(\tilde{\theta})=\frac{\theta^{2}/n}{\theta^{2}/(n-% 2)}=\frac{n-2}{n},

    which is less than one for every n≥3n\geq 3 and tends to one as n→∞n\to\infty: the estimator falls short of the bound for every finite sample size, but only by a factor which disappears in the limit.

  3. 3.
    ​

    The exponential distribution with rate θ\theta is an exponential family with h⁢(x)=1h(x)=1, η⁢(θ)=−θ\eta(\theta)=-\theta, T⁢(x)=xT(x)=x and B⁢(θ)=−log⁡θB(\theta)=-\log\theta, as in example 7.4, and thus μ⁢(θ)=𝔼θ⁢(X1)=1/θ\mu(\theta)=\mathbb{E}_{\theta}(X_{1})=1/\theta. By the discussion after proposition 13.5, an efficient estimator exists only for affine functions a+b/θa+b/\theta of μ⁢(θ)\mu(\theta). The rate g⁢(θ)=θg(\theta)=\theta is not of this form, and thus no unbiased estimator for θ\theta is efficient. The same conclusion can be reached directly: condition (13.2) with ℓ′⁢(θ)=n/θ−T\ell^{\prime}(\theta)=n/\theta-T would read W=θ+c⁢(θ)⁢n/θ−c⁢(θ)⁢TW=\theta+c(\theta)n/\theta-c(\theta)T, and since WW does not depend on θ\theta, the coefficients −c⁢(θ)-c(\theta) and θ+c⁢(θ)⁢n/θ\theta+c(\theta)n/\theta would have to be constants bb and aa, giving 𝔼θ⁢(W)=a+b⁢n/θ\mathbb{E}_{\theta}(W)=a+bn/\theta, which cannot equal θ\theta for all θ\theta.

  4. 4.
    ​

    The natural parameter η=−θ\eta=-\theta ranges over the open interval (−∞,0)(-\infty,0), and thus the family is of full rank and TT is complete and sufficient by corollary 8.3 and proposition 10.5. Since θ~\tilde{\theta} is a function of TT and unbiased for θ\theta, the Lehmann–Scheffe theorem, theorem 10.4, shows that it is the UMVUE for θ\theta. Together with the previous part this is an example of a UMVUE which is not efficient: the bound θ2/n\theta^{2}/n is not attained by any unbiased estimator, and the best unbiased estimator has variance θ2/(n−2)\theta^{2}/(n-2).

  5. 5.
    ​

    For g⁢(θ)=1/θg(\theta)=1/\theta we have g′⁢(θ)=−1/θ2g^{\prime}(\theta)=-1/\theta^{2}, and theorem 13.3 gives the bound

    g′⁢(θ)2ℐ︀n⁢(θ)=1/θ4n/θ2=1n⁢θ2=μ2n.\frac{g^{\prime}(\theta)^{2}}{\mathcal{I}_{n}(\theta)}=\frac{1/\theta^{4}}{n/% \theta^{2}}=\frac{1}{n\theta^{2}}=\frac{\mu^{2}}{n}.

    The sample mean is unbiased for μ\mu and has variance Varθ(X1)/n=μ2/n\mathop{\mathrm{Var}}\nolimits_{\theta}(X_{1})/n=\mu^{2}/n by lemma 2.2, and thus X¯\bar{X} is efficient for μ\mu, since the exchange identity (13.1) holds because the exponential distribution is a regular exponential family and X¯\bar{X} has finite variance, as noted after lemma 13.1. This is consistent with the previous parts: μ⁢(θ)=1/θ\mu(\theta)=1/\theta is the mean of the natural statistic, for which proposition 13.5 guarantees an efficient estimator, whereas the rate is a non-affine function of it. Whether an efficient estimator exists depends on which function of the parameter we set out to estimate.

∎

Solution to exercise 13.5.
  1. 1.
    ​

    Since μ\mu is given and τ=σ2\tau=\sigma^{2} is the parameter, the log-likelihood function is ℓ⁢(τ)=−n2⁢log⁡(2⁢π⁢τ)−12⁢τ⁢∑i(Xi−μ)2\ell(\tau)=-\tfrac{n}{2}\log(2\pi\tau)-\frac{1}{2\tau}\sum_{i}(X_{i}-\mu)^{2} and thus

    ℓ′⁢(τ)=−n2⁢τ+12⁢τ2⁢∑i=1n(Xi−μ)2,ℓ′′⁢(τ)=n2⁢τ2−1τ3⁢∑i=1n(Xi−μ)2.\ell^{\prime}(\tau)=-\frac{n}{2\tau}+\frac{1}{2\tau^{2}}\sum_{i=1}^{n}(X_{i}-% \mu)^{2},\qquad\ell^{\prime\prime}(\tau)=\frac{n}{2\tau^{2}}-\frac{1}{\tau^{3}% }\sum_{i=1}^{n}(X_{i}-\mu)^{2}.

    Since 𝔼⁢((Xi−μ)2)=τ\mathbb{E}\bigl{(}(X_{i}-\mu)^{2}\bigr{)}=\tau, we can use theorem 11.3 to find

    ℐ︀n⁢(τ)=−𝔼⁢(ℓ′′⁢(τ))=−n2⁢τ2+n⁢ττ3=n2⁢τ2=n2⁢σ4.\mathcal{I}_{n}(\tau)=-\mathbb{E}\bigl{(}\ell^{\prime\prime}(\tau)\bigr{)}=-% \frac{n}{2\tau^{2}}+\frac{n\tau}{\tau^{3}}=\frac{n}{2\tau^{2}}=\frac{n}{2% \sigma^{4}}.

    Thus, the Cramer–Rao bound is 2⁢σ4/n2\sigma^{4}/n.

  2. 2.
    ​

    The (Xi−μ)/σ(X_{i}-\mu)/\sigma are i.i.d. standard normally distributed, i.e. ∑i(Xi−μ)2/σ2∼χn2\sum_{i}(X_{i}-\mu)^{2}/\sigma^{2}\sim\chi^{2}_{n} by definition A.2, with mean nn and variance 2⁢n2n. Thus, σ^2=σ2n⁢∑i(Xi−μ)2/σ2\hat{\sigma}^{2}=\frac{\sigma^{2}}{n}\sum_{i}(X_{i}-\mu)^{2}/\sigma^{2} has mean σ2\sigma^{2} and variance σ4⋅2⁢n/n2=2⁢σ4/n\sigma^{4}\cdot 2n/n^{2}=2\sigma^{4}/n. This is the bound and thus σ^2\hat{\sigma}^{2} is efficient, since the exchange identity (13.1) holds because the normal distribution with known mean is a regular exponential family and σ^2\hat{\sigma}^{2} has finite variance, as noted after lemma 13.1. Alternatively, the score for the first part can be written as ℓ′⁢(σ2)=n2⁢σ4⁢(σ^2−σ2)\ell^{\prime}(\sigma^{2})=\frac{n}{2\sigma^{4}}(\hat{\sigma}^{2}-\sigma^{2}), which satisfies condition (13.2) with c⁢(σ2)=2⁢σ4/nc(\sigma^{2})=2\sigma^{4}/n.

  3. 3.
    ​

    From proposition A.3 we know that (n−1)⁢S2/σ2∼χn−12(n-1)S^{2}/\sigma^{2}\sim\chi^{2}_{n-1}, with variance 2⁢(n−1)2(n-1). Thus we have

    Var(S2)=σ4(n−1)2⋅2⁢(n−1)=2⁢σ4n−1>2⁢σ4n.\mathop{\mathrm{Var}}\nolimits(S^{2})=\frac{\sigma^{4}}{(n-1)^{2}}\cdot 2(n-1)% =\frac{2\sigma^{4}}{n-1}>\frac{2\sigma^{4}}{n}.

    Thus, the sample variance does not achieve the bound, and the efficiency of the sample variance is (n−1)/n(n-1)/n. Since the bound 2⁢σ4/n2\sigma^{4}/n still holds for unknown μ\mu, and since S2S^{2} is the UMVUE for σ2\sigma^{2} by example 10.7, this is a second example of a best unbiased estimator which is not efficient: no unbiased estimator for σ2\sigma^{2} can achieve the bound if μ\mu also needs to be estimated, and the loss of one degree of freedom is the price to be paid for not knowing μ\mu.

∎

Solution to exercise 13.6.
  1. 1.
    ​

    By proposition 11.4 the information of the sample is nn times the information of a single observation, in either parametrisation. Proposition 11.10 gives the relation for a single observation, and multiplying both sides by nn we find

    ℐ︀ψ⁢(ψ)=ℐ︀θ⁢(θ)g′⁢(θ)2for ⁢ψ=g⁢(θ).\mathcal{I}_{\psi}(\psi)=\frac{\mathcal{I}_{\theta}(\theta)}{g^{\prime}(\theta% )^{2}}\qquad\text{for }\psi=g(\theta).
  2. 2.
    ​

    Since ψ=g⁢(θ)\psi=g(\theta), the estimator WW is an unbiased estimator for the parameter ψ\psi, and theorem 13.2, applied in the ψ\psi-parametrisation, bounds its variance by 1/ℐ︀ψ⁢(ψ)1/\mathcal{I}_{\psi}(\psi). Using the first part we get

    1ℐ︀ψ⁢(ψ)=g′⁢(θ)2ℐ︀θ⁢(θ),\frac{1}{\mathcal{I}_{\psi}(\psi)}=\frac{g^{\prime}(\theta)^{2}}{\mathcal{I}_{% \theta}(\theta)},

    which is the bound of theorem 13.3. The two theorems thus give the same bound, and the bound does not depend on how the model is parametrised. This completes the proof.

∎

C.10 Multiparameter Fisher Information

Solution to exercise 14.1.
  1. 1.
    ​

    The Jacobian matrix of (μ,σ)↦(μ,σ2)(\mu,\sigma)\mapsto(\mu,\sigma^{2}) is given by

    J=(∂μ∂μ∂μ∂σ∂σ2∂μ∂σ2∂σ)=(1002⁢σ).J=\begin{pmatrix}\dfrac{\partial\mu}{\partial\mu}&\dfrac{\partial\mu}{\partial% \sigma}\\[9.42914pt] \dfrac{\partial\sigma^{2}}{\partial\mu}&\dfrac{\partial\sigma^{2}}{\partial% \sigma}\end{pmatrix}=\begin{pmatrix}1&0\\ 0&2\sigma\end{pmatrix}.

    This matrix is invertible for σ>0\sigma>0. Since all three matrices are diagonal, the product J⊤⁢ℐ︀n⁢(μ,σ2)⁢JJ^{\top}\mathcal{I}_{n}(\mu,\sigma^{2})J is also diagonal and consists of the entries 1⋅(n/σ2)⋅1=n/σ21\cdot(n/\sigma^{2})\cdot 1=n/\sigma^{2} and 2⁢σ⋅n/(2⁢σ4)⋅2⁢σ=2⁢n/σ22\sigma\cdot n/(2\sigma^{4})\cdot 2\sigma=2n/\sigma^{2}. Thus, by proposition 14.4, we have ℐ︀nφ⁢(μ,σ)=diag⁢(n/σ2, 2⁢n/σ2)\mathcal{I}^{\varphi}_{n}(\mu,\sigma)=\mathrm{diag}(n/\sigma^{2},\,2n/\sigma^{% 2}).

  2. 2.
    ​

    The log-likelihood function in terms of (μ,σ)(\mu,\sigma) is

    ℓ⁢(μ,σ)=−n2⁢log⁡(2⁢π)−n⁢log⁡σ−12⁢σ2⁢∑i=1n(Xi−μ)2.\ell(\mu,\sigma)=-\frac{n}{2}\log(2\pi)-n\log\sigma-\frac{1}{2\sigma^{2}}\sum_% {i=1}^{n}(X_{i}-\mu)^{2}.

    Taking derivatives we get ∂ℓ/∂μ=∑i(Xi−μ)/σ2\partial\ell/\partial\mu=\sum_{i}(X_{i}-\mu)/\sigma^{2} and ∂ℓ/∂σ=−n/σ+∑i(Xi−μ)2/σ3\partial\ell/\partial\sigma=-n/\sigma+\sum_{i}(X_{i}-\mu)^{2}/\sigma^{3}. Taking derivatives again, we find

    ∂2ℓ∂μ2\displaystyle\frac{\partial^{2}\ell}{\partial\mu^{2}}
    =−nσ2,\displaystyle=-\frac{n}{\sigma^{2}},
    ∂2ℓ∂μ⁢∂σ\displaystyle\frac{\partial^{2}\ell}{\partial\mu\,\partial\sigma}
    =−2σ3⁢∑i=1n(Xi−μ),\displaystyle=-\frac{2}{\sigma^{3}}\sum_{i=1}^{n}(X_{i}-\mu),
    ∂2ℓ∂σ2\displaystyle\frac{\partial^{2}\ell}{\partial\sigma^{2}}
    =nσ2−3σ4⁢∑i=1n(Xi−μ)2.\displaystyle=\frac{n}{\sigma^{2}}-\frac{3}{\sigma^{4}}\sum_{i=1}^{n}(X_{i}-% \mu)^{2}.

    Using 𝔼⁢(Xi−μ)=0\mathbb{E}(X_{i}-\mu)=0 and 𝔼⁢((Xi−μ)2)=σ2\mathbb{E}\bigl{(}(X_{i}-\mu)^{2}\bigr{)}=\sigma^{2} we find that the mixed derivative has expectation zero, and the last term has expectation n/σ2−3⁢n/σ2=−2⁢n/σ2n/\sigma^{2}-3n/\sigma^{2}=-2n/\sigma^{2}. Thus, by proposition 14.3, the information matrix is given by diag⁢(n/σ2, 2⁢n/σ2)\mathrm{diag}(n/\sigma^{2},\,2n/\sigma^{2}). This agrees with the result from part (a).

  3. 3.
    ​

    The inverse of the matrix from part (a) is given by diag⁢(σ2/n,σ2/(2⁢n))\mathrm{diag}\bigl{(}\sigma^{2}/n,\,\sigma^{2}/(2n)\bigr{)}. Thus, by theorem 14.5, any unbiased estimator for σ\sigma has variance of at least σ2/(2⁢n)\sigma^{2}/(2n). The sample variance S2S^{2} is an unbiased estimator for σ2\sigma^{2} by lemma 2.3, but the square root of the sample variance is not an unbiased estimator for σ\sigma: The function g⁢(x)=−xg(x)=-\sqrt{x} is strictly convex on (0,∞)(0,\infty) and by Jensen’s inequality, lemma A.7, we have −𝔼⁢(S2)≤𝔼⁢(−S2)-\sqrt{\mathbb{E}(S^{2})}\leq\mathbb{E}(-\sqrt{S^{2}}), i.e. 𝔼⁢(S)≤𝔼⁢(S2)=σ\mathbb{E}(S)\leq\sqrt{\mathbb{E}(S^{2})}=\sigma. The inequality is strict, since S2S^{2} is not constant with probability one (by proposition A.3 it is a multiple of a chi-squared random variable). Thus, 𝔼⁢(S)<σ\mathbb{E}(S)<\sigma, the estimator SS is biased for σ\sigma and the Cramer–Rao bound, which only concerns unbiased estimators, cannot be compared to Var(S)\mathop{\mathrm{Var}}\nolimits(S) directly.

∎

Solution to exercise 14.2.
  1. 1.
    ​

    We have ℓ1⁢(α,β)=α⁢log⁡β−log⁡Γ⁢(α)+(α−1)⁢log⁡X−β⁢X\ell_{1}(\alpha,\beta)=\alpha\log\beta-\log\Gamma(\alpha)+(\alpha-1)\log X-\beta X and

    ∂ℓ1∂α=log⁡β−ψ⁢(α)+log⁡Xand∂ℓ1∂β=αβ−X,\frac{\partial\ell_{1}}{\partial\alpha}=\log\beta-\psi(\alpha)+\log X\qquad% \text{and}\qquad\frac{\partial\ell_{1}}{\partial\beta}=\frac{\alpha}{\beta}-X,

    where ψ=(log⁡Γ)′\psi=(\log\Gamma)^{\prime} is the derivative of log⁡Γ\log\Gamma, which is smooth and thus differentiable on (0,∞)(0,\infty). Taking derivatives again, we get ∂2ℓ1/∂α2=−ψ′⁢(α)\partial^{2}\ell_{1}/\partial\alpha^{2}=-\psi^{\prime}(\alpha), ∂2ℓ1/∂α⁢∂β=1/β\partial^{2}\ell_{1}/\partial\alpha\,\partial\beta=1/\beta and ∂2ℓ1/∂β2=−α/β2\partial^{2}\ell_{1}/\partial\beta^{2}=-\alpha/\beta^{2}, as given in section 14.5. To verify the first identity of proposition 14.3, we note that the second component of the score has mean α/β−𝔼⁢(X)=0\alpha/\beta-\mathbb{E}(X)=0, since 𝔼⁢(X)=α/β\mathbb{E}(X)=\alpha/\beta by table A.2.

  2. 2.
    ​

    If α=2\alpha=2 is known, the model consists of one parameter β\beta and by theorem 11.3 the Fisher information is ℐ︀⁢(β)=−𝔼⁢(∂2ℓ1/∂β2)=α/β2=2/β2\mathcal{I}(\beta)=-\mathbb{E}\bigl{(}\partial^{2}\ell_{1}/\partial\beta^{2}% \bigr{)}=\alpha/\beta^{2}=2/\beta^{2}. The lower right corner of ℐ︀⁢(α,β)\mathcal{I}(\alpha,\beta) is the diagonal element ℐ︀⁢(α,β)22\mathcal{I}(\alpha,\beta)_{22}, which by definition is the variance of ∂ℓ1/∂β\partial\ell_{1}/\partial\beta with α\alpha fixed to the true value, i.e. for known shape. The diagonal element of the inverse answers a different question, asking about the estimation of β\beta when α\alpha is unknown. Part (d) shows that this value is larger.

  3. 3.
    ​

    The scale is θ=g⁢(β)=1/β\theta=g(\beta)=1/\beta where g′⁢(β)=−1/β2g^{\prime}(\beta)=-1/\beta^{2} and by proposition 11.10 we have

    ℐ︀θ⁢(θ)=ℐ︀β⁢(β)g′⁢(β)2=2/β21/β4=2⁢β2=2θ2.\mathcal{I}_{\theta}(\theta)=\frac{\mathcal{I}_{\beta}(\beta)}{g^{\prime}(% \beta)^{2}}=\frac{2/\beta^{2}}{1/\beta^{4}}=2\beta^{2}=\frac{2}{\theta^{2}}.

    This is the same value as in exercise 11.2.

  4. 4.
    ​

    We have a=ψ′⁢(α)a=\psi^{\prime}(\alpha), b=−1/βb=-1/\beta and c=α/β2c=\alpha/\beta^{2} and thus the determinant of ℐ︀⁢(α,β)\mathcal{I}(\alpha,\beta) is a⁢c−b2=(α⁢ψ′⁢(α)−1)/β2ac-b^{2}=\bigl{(}\alpha\psi^{\prime}(\alpha)-1\bigr{)}/\beta^{2}. The lower right element of the inverse is given by

    (ℐ︀⁢(α,β)−1)22=aa⁢c−b2=ψ′⁢(α)⁢β2α⁢ψ′⁢(α)−1.\bigl{(}\mathcal{I}(\alpha,\beta)^{-1}\bigr{)}_{22}=\frac{a}{ac-b^{2}}=\frac{% \psi^{\prime}(\alpha)\,\beta^{2}}{\alpha\psi^{\prime}(\alpha)-1}.

    For a sample of size nn, the information matrix is n⁢ℐ︀⁢(α,β)n\mathcal{I}(\alpha,\beta), the inverse is ℐ︀⁢(α,β)−1/n\mathcal{I}(\alpha,\beta)^{-1}/n and by theorem 14.5 we find the bound β2⁢ψ′⁢(α)/(n⁢(α⁢ψ′⁢(α)−1))\beta^{2}\psi^{\prime}(\alpha)/\bigl{(}n(\alpha\psi^{\prime}(\alpha)-1)\bigr{)} for an unbiased estimator for β\beta as claimed. The bound for known shape is 1/ℐ︀n⁢(β)=β2/(n⁢α)1/\mathcal{I}_{n}(\beta)=\beta^{2}/(n\alpha) and the ratio of the two is

    β2⁢ψ′⁢(α)/(n⁢(α⁢ψ′⁢(α)−1))β2/(n⁢α)=α⁢ψ′⁢(α)α⁢ψ′⁢(α)−1=11−ρ2,\frac{\beta^{2}\psi^{\prime}(\alpha)/\bigl{(}n(\alpha\psi^{\prime}(\alpha)-1)% \bigr{)}}{\beta^{2}/(n\alpha)}=\frac{\alpha\psi^{\prime}(\alpha)}{\alpha\psi^{% \prime}(\alpha)-1}=\frac{1}{1-\rho^{2}},

    where ρ2=1/(α⁢ψ′⁢(α))\rho^{2}=1/\bigl{(}\alpha\psi^{\prime}(\alpha)\bigr{)} is as in (14.2). For α=2\alpha=2 we have ψ′⁢(2)=π2/6−1≈0.6449\psi^{\prime}(2)=\pi^{2}/6-1\approx 0.6449 and thus α⁢ψ′⁢(α)≈1.2899\alpha\psi^{\prime}(\alpha)\approx 1.2899 and the ratio is approximately 1.2899/0.2899≈4.451.2899/0.2899\approx 4.45: not knowing the shape increases the bound for the rate by a factor of approximately 4.454.45, as in section 14.5.

∎

Solution to exercise 14.3.
  1. 1.
    ​

    Since the two samples are independent, the log-likelihood is the sum of the two sample log-likelihoods:

    ℓ⁢(μ1,μ2)=−m+n2⁢log⁡(2⁢π⁢σ2)−12⁢σ2⁢∑i=1m(Xi−μ1)2−12⁢σ2⁢∑j=1n(Yj−μ2)2,\ell(\mu_{1},\mu_{2})=-\frac{m+n}{2}\log(2\pi\sigma^{2})-\frac{1}{2\sigma^{2}}% \sum_{i=1}^{m}(X_{i}-\mu_{1})^{2}-\frac{1}{2\sigma^{2}}\sum_{j=1}^{n}(Y_{j}-% \mu_{2})^{2},

    with partial derivatives ∂ℓ/∂μ1=∑i(Xi−μ1)/σ2\partial\ell/\partial\mu_{1}=\sum_{i}(X_{i}-\mu_{1})/\sigma^{2} and ∂ℓ/∂μ2=∑j(Yj−μ2)/σ2\partial\ell/\partial\mu_{2}=\sum_{j}(Y_{j}-\mu_{2})/\sigma^{2}. The second partial derivatives are ∂2ℓ/∂μ12=−m/σ2\partial^{2}\ell/\partial\mu_{1}^{2}=-m/\sigma^{2}, ∂2ℓ/∂μ22=−n/σ2\partial^{2}\ell/\partial\mu_{2}^{2}=-n/\sigma^{2} and ∂2ℓ/∂μ1⁢∂μ2=0\partial^{2}\ell/\partial\mu_{1}\,\partial\mu_{2}=0, since each first partial derivative involves only one of the two parameters. None of these second derivatives are random, and thus proposition 14.3 gives ℐ︀⁢(μ1,μ2)=diag⁢(m/σ2,n/σ2)\mathcal{I}(\mu_{1},\mu_{2})=\mathrm{diag}(m/\sigma^{2},\,n/\sigma^{2}), without any expectations.

  2. 2.
    ​

    The difference is μ1−μ2=a⊤⁢θ\mu_{1}-\mu_{2}=a^{\top}\theta with a=(1,−1)a=(1,-1) and the inverse information matrix is diag⁢(σ2/m,σ2/n)\mathrm{diag}(\sigma^{2}/m,\,\sigma^{2}/n). By theorem 14.5 any unbiased estimator for μ1−μ2\mu_{1}-\mu_{2} has variance at least

    a⊤⁢ℐ︀⁢(μ1,μ2)−1⁢a=σ2m+σ2n.a^{\top}\mathcal{I}(\mu_{1},\mu_{2})^{-1}a=\frac{\sigma^{2}}{m}+\frac{\sigma^{% 2}}{n}.

    The estimator X¯−Y¯\bar{X}-\bar{Y} is unbiased for μ1−μ2\mu_{1}-\mu_{2} by lemma 2.2 and, since the two samples are independent, the variance of this estimator is Var(X¯)+Var(Y¯)=σ2/m+σ2/n\mathop{\mathrm{Var}}\nolimits(\bar{X})+\mathop{\mathrm{Var}}\nolimits(\bar{Y}% )=\sigma^{2}/m+\sigma^{2}/n. This is the bound. Thus, X¯−Y¯\bar{X}-\bar{Y} is an efficient estimator for the difference of the means.

  3. 3.
    ​

    If we include σ2\sigma^{2} as an additional parameter, the first derivative is

    ∂ℓ∂σ2=−m+n2⁢σ2+12⁢σ4⁢(∑i=1m(Xi−μ1)2+∑j=1n(Yj−μ2)2).\frac{\partial\ell}{\partial\sigma^{2}}=-\frac{m+n}{2\sigma^{2}}+\frac{1}{2% \sigma^{4}}\Bigl{(}\sum_{i=1}^{m}(X_{i}-\mu_{1})^{2}+\sum_{j=1}^{n}(Y_{j}-\mu_% {2})^{2}\Bigr{)}.

    The mixed derivatives with respect to μ1\mu_{1} and σ2\sigma^{2} and with respect to μ2\mu_{2} and σ2\sigma^{2} are −∑i(Xi−μ1)/σ4-\sum_{i}(X_{i}-\mu_{1})/\sigma^{4} and −∑j(Yj−μ2)/σ4-\sum_{j}(Y_{j}-\mu_{2})/\sigma^{4}, both with expectation zero, and the second derivative with respect to σ2\sigma^{2} has expectation (m+n)/(2⁢σ4)−(m+n)⁢σ2/σ6=−(m+n)/(2⁢σ4)(m+n)/(2\sigma^{4})-(m+n)\sigma^{2}/\sigma^{6}=-(m+n)/(2\sigma^{4}), as shown in example 14.6. Thus, the information matrix is

    ℐ︀⁢(μ1,μ2,σ2)=diag⁢(mσ2,nσ2,m+n2⁢σ4),\mathcal{I}(\mu_{1},\mu_{2},\sigma^{2})=\mathrm{diag}\Bigl{(}\frac{m}{\sigma^{% 2}},\,\frac{n}{\sigma^{2}},\,\frac{m+n}{2\sigma^{4}}\Bigr{)},

    still diagonal, and thus σ2\sigma^{2} is orthogonal to both means. The upper left 2×22\times 2 block of the inverse is the inverse information matrix from part (b), and the Cramer–Rao bound for μ1−μ2\mu_{1}-\mu_{2} is still σ2/m+σ2/n\sigma^{2}/m+\sigma^{2}/n, and this bound is achieved by X¯−Y¯\bar{X}-\bar{Y}, whose definition does not depend on σ2\sigma^{2}. For point estimation of the difference, not knowing the variance does not change anything. However, once we want to determine the accuracy of an estimate, the unknown variance starts to matter, since the bound depends on σ2\sigma^{2} and thus needs to be estimated. This is the reason why we go from the normal distribution to the tt distribution in lectures 25 and 29.

∎

C.11 Consistency of the MLE

Solution to exercise 16.1.
  1. 1.
    ​

    For (R1) we have two different rates and thus two different distributions: for example, the means 1/θ1/\theta are different. For (R2) the support is [0,∞)[0,\infty) for all θ\theta. For (R3) we have log⁡f⁢(x;θ)=log⁡θ−θ⁢x\log f(x;\theta)=\log\theta-\theta x and the derivatives w.r.t. θ\theta are 1/θ−x1/\theta-x, −1/θ2-1/\theta^{2} and 2/θ32/\theta^{3}, all of which are continuous at θ\theta. The derivatives of ff are of the form (polynomial in xx and θ\theta) times e−θ⁢xe^{-\theta x} and thus are integrable over [0,∞)[0,\infty) uniformly for θ\theta in any compact subinterval of (0,∞)(0,\infty). We can therefore interchange the order of differentiation and integration, and ℐ︀⁢(θ)=1/θ2\mathcal{I}(\theta)=1/\theta^{2} is continuous and lies strictly between zero and infinity, so that (R3) holds. Finally, for (R4), the third derivative 2/θ32/\theta^{3} does not depend on xx and for |θ−θ0|<δ|\theta-\theta_{0}|<\delta with δ=θ0/2\delta=\theta_{0}/2 we have θ>θ0/2\theta>\theta_{0}/2 and thus 2/θ3≤16/θ032/\theta^{3}\leq 16/\theta_{0}^{3}. The constant M⁢(x)=16/θ03M(x)=16/\theta_{0}^{3} bounds the third derivative and has finite expectation.

  2. 2.
    ​

    The likelihood equation n/θ−∑ixi=0n/\theta-\sum_{i}x_{i}=0 has only one solution θ=1/x¯\theta=1/\bar{x}, since ∑ixi>0\sum_{i}x_{i}>0 with probability one. Since ℓ′′⁢(θ)=−n/θ2<0\ell^{\prime\prime}(\theta)=-n/\theta^{2}<0, the log-likelihood function is strictly concave and thus the root is the global maximum of the likelihood function, i.e. the MLE as in example 5.2. Since all conditions of theorem 16.3 are satisfied, we find that 1/X¯1/\bar{X} is consistent for θ\theta.

  3. 3.
    ​

    By the law of large numbers, theorem A.14, we have X¯→p𝔼θ⁢(X1)=1/θ\bar{X}\xrightarrow{\ \mathrm{p}\ }\mathbb{E}_{\theta}(X_{1})=1/\theta. Since the function g⁢(y)=1/yg(y)=1/y is continuous at y=1/θ>0y=1/\theta>0, the continuous mapping theorem, theorem A.17, applies and we find 1/X¯→pg⁢(1/θ)=θ1/\bar{X}\xrightarrow{\ \mathrm{p}\ }g(1/\theta)=\theta. This proves the consistency of the MLE.

∎

Solution to exercise 16.2.
  1. 1.
    ​

    Since f⁢(x;θ)=θx⁢(1−θ)1−xf(x;\theta)=\theta^{x}(1-\theta)^{1-x} for x∈{0,1}x\in\{0,1\}, we have

    log⁡f⁢(X1;θ)f⁢(X1;θ0)=X1⁢log⁡θθ0+(1−X1)⁢log⁡1−θ1−θ0,\log\frac{f(X_{1};\theta)}{f(X_{1};\theta_{0})}=X_{1}\log\frac{\theta}{\theta_% {0}}+(1-X_{1})\log\frac{1-\theta}{1-\theta_{0}},

    and taking expectations with 𝔼θ0⁢(X1)=θ0\mathbb{E}_{\theta_{0}}(X_{1})=\theta_{0} gives

    D⁢(θ)=θ0⁢log⁡θθ0+(1−θ0)⁢log⁡1−θ1−θ0.D(\theta)=\theta_{0}\log\frac{\theta}{\theta_{0}}+(1-\theta_{0})\log\frac{1-% \theta}{1-\theta_{0}}.
  2. 2.
    ​

    Using the rule log⁡y≤y−1\log y\leq y-1 for logarithms, we find

    D⁢(θ)≤θ0⁢(θθ0−1)+(1−θ0)⁢(1−θ1−θ0−1)=(θ−θ0)+(θ0−θ)=0.D(\theta)\leq\theta_{0}\Bigl{(}\frac{\theta}{\theta_{0}}-1\Bigr{)}+(1-\theta_{% 0})\Bigl{(}\frac{1-\theta}{1-\theta_{0}}-1\Bigr{)}=(\theta-\theta_{0})+(\theta% _{0}-\theta)=0.

    For D⁢(θ)=0D(\theta)=0 to hold, we need to have equality in both applications of the inequality, i.e. we need θ/θ0=1\theta/\theta_{0}=1 and (1−θ)/(1−θ0)=1(1-\theta)/(1-\theta_{0})=1. Both of these conditions state θ=θ0\theta=\theta_{0}. Thus we have D⁢(θ)<0D(\theta)<0 for all θ≠θ0\theta\neq\theta_{0}, as stated in theorem 16.2.

∎

Solution to exercise 16.3.
  1. 1.
    ​

    Since all observations are at most θ\theta, we have θ^n≤θ\hat{\theta}_{n}\leq\theta and thus |θ^n−θ|>ε|\hat{\theta}_{n}-\theta|>\varepsilon is the same as θ^n<θ−ε\hat{\theta}_{n}<\theta-\varepsilon. This is the case if and only if all nn observations are less than θ−ε\theta-\varepsilon. Using independence we find

    ℙθ⁢(|θ^n−θ|>ε)=(θ−εθ)n=(1−εθ)n⟶0(n→∞),\mathbb{P}_{\theta}\bigl{(}|\hat{\theta}_{n}-\theta|>\varepsilon\bigr{)}=\Bigl% {(}\frac{\theta-\varepsilon}{\theta}\Bigr{)}^{n}=\Bigl{(}1-\frac{\varepsilon}{% \theta}\Bigr{)}^{n}\longrightarrow 0\qquad(n\to\infty),

    since 0<1−ε/θ<10<1-\varepsilon/\theta<1. For ε≥θ\varepsilon\geq\theta, the probability of this event is zero. Thus θ^n→pθ\hat{\theta}_{n}\xrightarrow{\ \mathrm{p}\ }\theta and the MLE is consistent.

  2. 2.
    ​

    Condition (R2) is violated, since the support (0,θ)(0,\theta) depends on θ\theta and (R3) is also violated: for fixed xx, the function θ↦f⁢(x;θ)\theta\mapsto f(x;\theta) has a jump from 0 to 1/θ1/\theta at θ=x\theta=x and thus is not differentiable there. The uniqueness of the MLE fails, since the likelihood equation has no roots: on the region θ≥maxi⁡xi\theta\geq\max_{i}x_{i} where LL is positive, the derivative −n⁢θ−n−1-n\theta^{-n-1} never equals zero. Since the proof of theorem 16.3 requires a differentiable likelihood and a root of the derivative to apply Rolle’s theorem, neither of these conditions are satisfied. Nevertheless, the statement is true, as can be seen by the direct computation in the first part of the question, or using the mean squared error from exercise 2.2.

∎

C.12 Asymptotic Normality of the MLE

Solution to exercise 17.1.
  1. 1.
    ​

    Since we have ℐ︀⁢(θ)=1/(θ⁢(1−θ))\mathcal{I}(\theta)=1/\bigl{(}\theta(1-\theta)\bigr{)}, theorem 17.1 gives that n⁢(X¯−θ)→dN⁢(0,θ⁢(1−θ))\sqrt{n}\,(\bar{X}-\theta)\xrightarrow{\ \mathrm{d}\ }N\bigl{(}0,\theta(1-% \theta)\bigr{)}. Since the mean and variance of the Bernoulli distribution equal θ\theta and θ⁢(1−θ)\theta(1-\theta), respectively, this is the central limit theorem, theorem A.15, for X¯\bar{X} and thus the two results coincide.

  2. 2.
    ​

    From theorem 5.6 we know that the MLE for ψ\psi is given by ψ^n=log⁡(X¯/(1−X¯))\hat{\psi}_{n}=\log\bigl{(}\bar{X}/(1-\bar{X})\bigr{)}. Taking derivatives we find the derivative of g⁢(θ)=log⁡θ−log⁡(1−θ)g(\theta)=\log\theta-\log(1-\theta) to be g′⁢(θ)=1/θ+1/(1−θ)=1/(θ⁢(1−θ))g^{\prime}(\theta)=1/\theta+1/(1-\theta)=1/\bigl{(}\theta(1-\theta)\bigr{)} and using the delta method, theorem A.18, we find

    n⁢(ψ^n−ψ)→dN⁢(0,θ⁢(1−θ)θ2⁢(1−θ)2)=N⁢(0,1θ⁢(1−θ)).\sqrt{n}\,(\hat{\psi}_{n}-\psi)\xrightarrow{\ \mathrm{d}\ }N\Bigl{(}0,\ \frac{% \theta(1-\theta)}{\theta^{2}(1-\theta)^{2}}\Bigr{)}=N\Bigl{(}0,\frac{1}{\theta% (1-\theta)}\Bigr{)}.
  3. 3.
    ​

    Substituting θ\theta by θ^n\hat{\theta}_{n} in the asymptotic variance, we find the standard error se(ψ^n)=1/n⁢θ^n⁢(1−θ^n)\mathop{\mathrm{se}}\nolimits(\hat{\psi}_{n})=1/\sqrt{n\hat{\theta}_{n}(1-\hat% {\theta}_{n})} and thus the approximate 95%95\% interval ψ^n±1.96/n⁢θ^n⁢(1−θ^n)\hat{\psi}_{n}\pm 1.96/\sqrt{n\hat{\theta}_{n}(1-\hat{\theta}_{n})}. Both of these quantities are undefined if all observations are equal, i.e. at the boundary of the domain discussed in section 17.5.

∎

Solution to exercise 17.2.
  1. 1.
    ​

    Using the result ℐ︀⁢(θ)=1/θ\mathcal{I}(\theta)=1/\theta from the theorem, we find n⁢(X¯−θ)→dN⁢(0,θ)\sqrt{n}\,(\bar{X}-\theta)\xrightarrow{\ \mathrm{d}\ }N(0,\theta). Since the Poisson distribution has mean and variance equal to θ\theta, this is the central limit theorem for X¯\bar{X}.

  2. 2.
    ​

    For g⁢(θ)=e−θg(\theta)=e^{-\theta} we have g′⁢(θ)=−e−θg^{\prime}(\theta)=-e^{-\theta} and using the delta method we find

    n⁢(p^0−p0)→dN⁢(0,e−2⁢θ⁢θ).\sqrt{n}\,(\hat{p}_{0}-p_{0})\xrightarrow{\ \mathrm{d}\ }N\bigl{(}0,\ e^{-2% \theta}\theta\bigr{)}.

    In exercise 5.7 we found 𝔼θ⁢(p^0)=exp⁡(n⁢θ⁢(e−1/n−1))\mathbb{E}_{\theta}(\hat{p}_{0})=\exp\bigl{(}n\theta(e^{-1/n}-1)\bigr{)} and expanding e−1/n=1−1/n+1/(2⁢n2)−⋯e^{-1/n}=1-1/n+1/(2n^{2})-\dotsb we find that the bias is approximately θ⁢e−θ/(2⁢n)\theta e^{-\theta}/(2n), which is of order 1/n1/n. Since it is multiplied by n\sqrt{n}, the bias goes to zero and thus is of a smaller order than the spread 1/n1/\sqrt{n} and does not contribute in the limit: the estimator is biased for finite nn, but is asymptotically normal with mean p0p_{0}.

∎

Solution to exercise 17.3.
  1. 1.
    ​

    For x≥0x\geq 0 and n>x/θn>x/\theta, the event n⁢(θ−θ^n)>xn(\theta-\hat{\theta}_{n})>x is the event θ^n<θ−x/n\hat{\theta}_{n}<\theta-x/n, i.e. that all observations are less than θ−x/n\theta-x/n. Thus we have

    ℙθ⁢(n⁢(θ−θ^n)>x)=(1−xn⁢θ)n⟶e−x/θ(n→∞),\mathbb{P}_{\theta}\bigl{(}n(\theta-\hat{\theta}_{n})>x\bigr{)}=\Bigl{(}1-% \frac{x}{n\theta}\Bigr{)}^{n}\longrightarrow e^{-x/\theta}\qquad(n\to\infty),

    where we used the standard limit (1−a/n)n→e−a(1-a/n)^{n}\to e^{-a}. The right-hand side is the probability that an exponential random variable with mean θ\theta is greater than xx. Thus, n⁢(θ−θ^n)n(\theta-\hat{\theta}_{n}) converges in distribution to the corresponding exponential distribution.

  2. 2.
    ​

    Since n⁢(θ^n−θ)=−n−1/2⋅n⁢(θ−θ^n)\sqrt{n}\,(\hat{\theta}_{n}-\theta)=-n^{-1/2}\cdot n(\theta-\hat{\theta}_{n}), the second factor converges in distribution to the result from the first part, and the first factor is a constant which converges to zero. Thus, by Slutsky’s theorem (theorem A.16), the product converges in distribution to 0⋅Y=00\cdot Y=0. This is equivalent to convergence in probability to zero. Thus we cannot have a limit N⁢(0,τ2)N(0,\tau^{2}) where τ2>0\tau^{2}>0, since the error of the MLE is of order 1/n1/n and not of order 1/n1/\sqrt{n}. This is a manifestation of the super-efficiency from lecture 13, seen through the lens of the limiting distribution: the variance of the bias-corrected maximum is of order 1/n21/n^{2}, smaller than allowed by the Cramer–Rao bound, and thus neither the bound nor theorem 17.1 apply, since the model violates (R2).

∎

C.13 The EM Algorithm

Solution to exercise 19.1.
  1. 1.
    ​

    Using the formula for the sum of a geometric series, we find

    ℙp⁢(X≥m)\displaystyle\mathbb{P}_{p}(X\geq m)
    =∑x=m∞(1−p)x−1⁢p=p⁢(1−p)m−1⁢∑j=0∞(1−p)j\displaystyle=\sum_{x=m}^{\infty}(1-p)^{x-1}p=p\,(1-p)^{m-1}\sum_{j=0}^{\infty% }(1-p)^{j}
    =p⁢(1−p)m−1⋅1p=(1−p)m−1.\displaystyle=p\,(1-p)^{m-1}\cdot\frac{1}{p}=(1-p)^{m-1}.

    Since the event {X≥m+k}\{X\geq m+k\} is contained in the event {X≥m}\{X\geq m\}, we can use the definition of conditional probability to get

    ℙp⁢(X≥m+k|X≥m)=(1−p)m+k−1(1−p)m−1=(1−p)k=ℙp⁢(X≥k+1).\mathbb{P}_{p}(X\geq m+k\mskip 1.0mu|\mskip 1.0muX\geq m)=\frac{(1-p)^{m+k-1}}% {(1-p)^{m-1}}=(1-p)^{k}=\mathbb{P}_{p}(X\geq k+1).

    This shows that, given that the first m−1m-1 trials were failures, the number X−(m−1)X-(m-1) of additional trials required until the first success is still geometrically distributed. The information about the number of failures so far does not change the distribution of the number of additional failures required. This is the discrete analogue of the statement ℙθ⁢(X>s+t⁢|X>⁢s)=ℙθ⁢(X>t)\mathbb{P}_{\theta}(X>s+t\mskip 1.0mu|\mskip 1.0muX>s)=\mathbb{P}_{\theta}(X>t) for the exponential distribution in exercise 5.6. The shift by one is due to the fact that the support of the geometric distribution starts at 11 instead of at 0.

  2. 2.
    ​

    Given X≥mX\geq m, the conditional probability weights of XX are

    f⁢(x;p)ℙp⁢(X≥m)=(1−p)x−m⁢p\frac{f(x;p)}{\mathbb{P}_{p}(X\geq m)}=(1-p)^{x-m}p

    for x≥mx\geq m. Substituting x=m−1+jx=m-1+j with j∈{1,2,…}j\in\{1,2,\dots\} we get

    𝔼p⁢(X|X≥m)\displaystyle\mathbb{E}_{p}(X\mskip 1.0mu|\mskip 1.0muX\geq m)
    =∑j=1∞(m−1+j)⁢(1−p)j−1⁢p\displaystyle=\sum_{j=1}^{\infty}(m-1+j)(1-p)^{j-1}p
    =(m−1)⁢∑j=1∞(1−p)j−1⁢p+∑j=1∞j⁢(1−p)j−1⁢p\displaystyle=(m-1)\sum_{j=1}^{\infty}(1-p)^{j-1}p+\sum_{j=1}^{\infty}j(1-p)^{% j-1}p
    =(m−1)+𝔼p⁢(X),\displaystyle=(m-1)+\mathbb{E}_{p}(X),

    where the first sum is the total probability 11 and the second sum is the mean of the geometric distribution. Using 𝔼p⁢(X)=1/p\mathbb{E}_{p}(X)=1/p we get the claim. Alternatively we can get the same result from the first part of the solution: given X≥mX\geq m, the value X−(m−1)X-(m-1) is geometrically distributed with parameter pp and thus has mean 1/p1/p.

∎

Solution to exercise 19.2.
  1. 1.
    ​

    The log-likelihood function for the given data is

    ℓ⁢(π)=∑i=1nlog⁡(π⁢f1⁢(xi)+(1−π)⁢f2⁢(xi)).\ell(\pi)=\sum_{i=1}^{n}\log\bigl{(}\pi f_{1}(x_{i})+(1-\pi)f_{2}(x_{i})\bigr{% )}.

    Taking derivatives we get the likelihood equation

    ℓ′⁢(π)=∑i=1nf1⁢(xi)−f2⁢(xi)π⁢f1⁢(xi)+(1−π)⁢f2⁢(xi)=0.\ell^{\prime}(\pi)=\sum_{i=1}^{n}\frac{f_{1}(x_{i})-f_{2}(x_{i})}{\pi f_{1}(x_% {i})+(1-\pi)f_{2}(x_{i})}=0.

    This equation cannot be solved analytically for general densities.

  2. 2.
    ​

    The density of the complete data is fc⁢(x,1;π)=π⁢f1⁢(x)f_{c}(x,1;\pi)=\pi f_{1}(x) and fc⁢(x,2;π)=(1−π)⁢f2⁢(x)f_{c}(x,2;\pi)=(1-\pi)f_{2}(x). Thus we have

    ℓc⁢(π)=∑i=1n[𝟏{Zi=1}⁢(log⁡π+log⁡f1⁢(xi))+𝟏{Zi=2}⁢(log⁡(1−π)+log⁡f2⁢(xi))].\ell_{c}(\pi)=\sum_{i=1}^{n}\Bigl{[}\mathbf{1}_{\{Z_{i}=1\}}\bigl{(}\log\pi+% \log f_{1}(x_{i})\bigr{)}+\mathbf{1}_{\{Z_{i}=2\}}\bigl{(}\log(1-\pi)+\log f_{% 2}(x_{i})\bigr{)}\Bigr{]}.

    Using Bayes’ rule we find the conditional probability of Zi=1Z_{i}=1 given Xi=xiX_{i}=x_{i} under πk\pi_{k} as γi⁢(πk)=πk⁢f1⁢(xi)/(πk⁢f1⁢(xi)+(1−πk)⁢f2⁢(xi))\gamma_{i}(\pi_{k})=\pi_{k}f_{1}(x_{i})/\bigl{(}\pi_{k}f_{1}(x_{i})+(1-\pi_{k}% )f_{2}(x_{i})\bigr{)}. Substituting the indicators using these responsibilities we get

    Q⁢(π|πk)=log⁡π⁢∑i=1nγi⁢(πk)+log⁡(1−π)⁢∑i=1n(1−γi⁢(πk))+const,Q(\pi\mskip 1.0mu|\mskip 1.0mu\pi_{k})=\log\pi\sum_{i=1}^{n}\gamma_{i}(\pi_{k}% )+\log(1-\pi)\sum_{i=1}^{n}\bigl{(}1-\gamma_{i}(\pi_{k})\bigr{)}+\text{const},

    where the constant term combines the terms log⁡f1⁢(xi)\log f_{1}(x_{i}) and log⁡f2⁢(xi)\log f_{2}(x_{i}), which do not depend on π\pi. Setting the derivative ∑iγi/π−∑i(1−γi)/(1−π)\sum_{i}\gamma_{i}/\pi-\sum_{i}(1-\gamma_{i})/(1-\pi) equal to zero we find πk+1=1n⁢∑iγi⁢(πk)\pi_{k+1}=\frac{1}{n}\sum_{i}\gamma_{i}(\pi_{k}) and the second derivative −∑iγi/π2−∑i(1−γi)/(1−π)2-\sum_{i}\gamma_{i}/\pi^{2}-\sum_{i}(1-\gamma_{i})/(1-\pi)^{2} is negative for every π∈(0,1)\pi\in(0,1), so Q⁢(⋅|πk)Q(\,\mathord{\cdot}\,\mskip 1.0mu|\mskip 1.0mu\pi_{k}) is concave on (0,1)(0,1) and the stationary point is the global maximum.

  3. 3.
    ​

    Using the notation fi=f⁢(xi;π)=π⁢f1⁢(xi)+(1−π)⁢f2⁢(xi)f_{i}=f(x_{i};\pi)=\pi f_{1}(x_{i})+(1-\pi)f_{2}(x_{i}) for the denominator, we have

    γi⁢(π)−π\displaystyle\gamma_{i}(\pi)-\pi
    =π⁢f1⁢(xi)−π⁢(π⁢f1⁢(xi)+(1−π)⁢f2⁢(xi))fi\displaystyle=\frac{\pi f_{1}(x_{i})-\pi\bigl{(}\pi f_{1}(x_{i})+(1-\pi)f_{2}(% x_{i})\bigr{)}}{f_{i}}
    =π⁢(1−π)⁢(f1⁢(xi)−f2⁢(xi))fi.\displaystyle=\frac{\pi(1-\pi)\bigl{(}f_{1}(x_{i})-f_{2}(x_{i})\bigr{)}}{f_{i}}.

    Summing over ii and comparing to the first part of the solution we find 1n⁢∑iγi⁢(π)−π=π⁢(1−π)⁢ℓ′⁢(π)/n\frac{1}{n}\sum_{i}\gamma_{i}(\pi)-\pi=\pi(1-\pi)\,\ell^{\prime}(\pi)/n. A value π∗∈(0,1)\pi^{*}\in(0,1) is a fixed point of the iteration if and only if the left-hand side is zero at π∗\pi^{*} and since π∗⁢(1−π∗)>0\pi^{*}(1-\pi^{*})>0 this is the case if and only if ℓ′⁢(π∗)=0\ell^{\prime}(\pi^{*})=0. Thus the fixed points of EM are the roots of the likelihood equation as given in proposition 19.3.

∎

Solution to exercise 19.3.

At θ0=(1/2,2,4)\theta_{0}=(1/2,2,4), the mixture density of observation ii is 12⁢φ1⁢(xi−2)+12⁢φ1⁢(xi−4)\frac{1}{2}\varphi_{1}(x_{i}-2)+\frac{1}{2}\varphi_{1}(x_{i}-4). Using the standard normal density values φ1⁢(1)=0.2420\varphi_{1}(1)=0.2420, φ1⁢(2)=0.05399\varphi_{1}(2)=0.05399, φ1⁢(3)=0.004432\varphi_{1}(3)=0.004432, φ1⁢(4)=0.0001338\varphi_{1}(4)=0.0001338 and φ1⁢(5)=0.0000015\varphi_{1}(5)=0.0000015, we find the five densities to be 0.0270620.027062, 0.123200.12320, 0.123200.12320, 0.0270620.027062 and 0.00221670.0022167. The corresponding logarithms are

−3.6096,−2.0939,−2.0939,−3.6096,−6.1118-3.6096,\quad-2.0939,\quad-2.0939,\quad-3.6096,\quad-6.1118

and thus ℓ⁢(θ0)=−17.5188\ell(\theta_{0})=-17.5188. At θ1=(0.4001,0.5445,5.9709)\theta_{1}=(0.4001,0.5445,5.9709) the densities are 0.137630.13763, 0.143890.14389, 0.149390.14939, 0.239220.23922 and 0.140940.14094. The corresponding logarithms are

−1.9832,−1.9387,−1.9012,−1.4304,−1.9595-1.9832,\quad-1.9387,\quad-1.9012,\quad-1.4304,\quad-1.9595

and thus ℓ⁢(θ1)=−9.2129\ell(\theta_{1})=-9.2129. The log-likelihood has increased by 8.318.31 as required by theorem 19.2. The largest contribution, 4.154.15, comes from observation x5=7x_{5}=7: under θ0\theta_{0} this observation is three standard deviations away from the nearer mean 44 and has tiny density, whereas under θ1\theta_{1} it is about one standard deviation away from the mean 5.975.97. The observations x1=0x_{1}=0 and x4=6x_{4}=6 contribute 1.631.63 and 2.182.18, respectively, for the same reason. The observations x2=1x_{2}=1 and x3=5x_{3}=5, which were always within one standard deviation of a mean, contribute little. ∎

Solution to exercise 19.4.
  1. 1.
    ​

    Since we know π\pi, the complete-data log-likelihood is the same as in example 19.4 and the terms log⁡π\log\pi and log⁡(1−π)\log(1-\pi) are now constants. The E step determines the responsibilities γi=π⁢φσ⁢(xi−μ1,k)/(π⁢φσ⁢(xi−μ1,k)+(1−π)⁢φσ⁢(xi−μ2,k))\gamma_{i}=\pi\varphi_{\sigma}(x_{i}-\mu_{1,k})/\bigl{(}\pi\varphi_{\sigma}(x_% {i}-\mu_{1,k})+(1-\pi)\varphi_{\sigma}(x_{i}-\mu_{2,k})\bigr{)}. Note that Q⁢(θ|θk)Q(\theta\mskip 1.0mu|\mskip 1.0mu\theta_{k}) depends on θ=(μ1,μ2)\theta=(\mu_{1},\mu_{2}) only via the terms −12⁢σ2⁢∑i[γi⁢(xi−μ1)2+(1−γi)⁢(xi−μ2)2]-\frac{1}{2\sigma^{2}}\sum_{i}\bigl{[}\gamma_{i}(x_{i}-\mu_{1})^{2}+(1-\gamma_% {i})(x_{i}-\mu_{2})^{2}\bigr{]}. The partial derivatives are as shown in the example, and setting them equal to zero gives the two mean updates of (19.2). The update for the weight is no longer needed.

  2. 2.
    ​

    Since φσ′⁢(u)=−(u/σ2)⁢φσ⁢(u)\varphi_{\sigma}^{\prime}(u)=-(u/\sigma^{2})\varphi_{\sigma}(u), we have ∂∂μ1⁢φσ⁢(x−μ1)=x−μ1σ2⁢φσ⁢(x−μ1)\frac{\partial}{\partial\mu_{1}}\varphi_{\sigma}(x-\mu_{1})=\frac{x-\mu_{1}}{% \sigma^{2}}\varphi_{\sigma}(x-\mu_{1}) and thus we can take derivatives of ℓ⁢(θ)=∑ilog⁡(π⁢φσ⁢(xi−μ1)+(1−π)⁢φσ⁢(xi−μ2))\ell(\theta)=\sum_{i}\log\bigl{(}\pi\varphi_{\sigma}(x_{i}-\mu_{1})+(1-\pi)% \varphi_{\sigma}(x_{i}-\mu_{2})\bigr{)} to get

    ∂ℓ∂μ1⁢(θ)\displaystyle\frac{\partial\ell}{\partial\mu_{1}}(\theta)
    =∑i=1nπ⁢φσ⁢(xi−μ1)π⁢φσ⁢(xi−μ1)+(1−π)⁢φσ⁢(xi−μ2)⋅xi−μ1σ2\displaystyle=\sum_{i=1}^{n}\frac{\pi\varphi_{\sigma}(x_{i}-\mu_{1})}{\pi% \varphi_{\sigma}(x_{i}-\mu_{1})+(1-\pi)\varphi_{\sigma}(x_{i}-\mu_{2})}\cdot% \frac{x_{i}-\mu_{1}}{\sigma^{2}}
    =1σ2⁢∑i=1nγi⁢(θ)⁢(xi−μ1),\displaystyle=\frac{1}{\sigma^{2}}\sum_{i=1}^{n}\gamma_{i}(\theta)\,(x_{i}-\mu% _{1}),

    and similarly we find ∂ℓ∂μ2⁢(θ)=1σ2⁢∑i(1−γi⁢(θ))⁢(xi−μ2)\frac{\partial\ell}{\partial\mu_{2}}(\theta)=\frac{1}{\sigma^{2}}\sum_{i}\bigl% {(}1-\gamma_{i}(\theta)\bigr{)}(x_{i}-\mu_{2}). Setting both of these equal to zero gives μ1=∑iγi⁢(θ)⁢xi/∑iγi⁢(θ)\mu_{1}=\sum_{i}\gamma_{i}(\theta)x_{i}/\sum_{i}\gamma_{i}(\theta) and the corresponding equation for μ2\mu_{2}, stating that θ\theta is mapped to itself by the updates of the first part. Thus, the likelihood equations and the fixed-point equations of the EM algorithm coincide as predicted by proposition 19.3, and we also find the identity of this proposition directly: the derivative of ℓ\ell equals the derivative of θ′↦Q⁢(θ′|θ)\theta^{\prime}\mapsto Q(\theta^{\prime}\mskip 1.0mu|\mskip 1.0mu\theta) at θ′=θ\theta^{\prime}=\theta.

  3. 3.
    ​

    As π→1\pi\to 1, the responsibilities converge to γi=1\gamma_{i}=1 for all observations since φσ⁢(xi−μ1,k)>0\varphi_{\sigma}(x_{i}-\mu_{1,k})>0. Thus, the update for μ1\mu_{1} is now μ1,k+1=x¯\mu_{1,k+1}=\bar{x} independent of the initial value, and this is the MLE for the mean of a single normal sample in example 5.4, reached in one step. The update for μ2\mu_{2} is now the undefined ratio 0/00/0: a component with no weight has no observations to use to estimate its mean, and the mean of this component is not identifiable as discussed in section 19.5.

∎

C.14 Hypothesis Testing

Solution to exercise 20.1.
  1. 1.
    ​

    By equation (20.2) the power at θ1\theta_{1} is β⁢(θ1)=Φ⁢(d−z1−α)\beta(\theta_{1})=\Phi(d-z_{1-\alpha}), where d=n⁢(θ1−θ0)/σd=\sqrt{n}\,(\theta_{1}-\theta_{0})/\sigma. Since Φ\Phi is strictly increasing with Φ⁢(z1−γ)=1−γ\Phi(z_{1-\gamma})=1-\gamma, we have β⁢(θ1)≥1−γ\beta(\theta_{1})\geq 1-\gamma if and only if d−z1−α≥z1−γd-z_{1-\alpha}\geq z_{1-\gamma}, i.e. if and only if

    n≥(z1−α+z1−γ)⁢σθ1−θ0.\sqrt{n}\geq\frac{(z_{1-\alpha}+z_{1-\gamma})\,\sigma}{\theta_{1}-\theta_{0}}.

    Since 1−γ>α1-\gamma>\alpha we have z1−γ>zα=−z1−αz_{1-\gamma}>z_{\alpha}=-z_{1-\alpha}, and thus the right-hand side is positive; squaring both sides gives the stated condition.

  2. 2.
    ​

    With z0.95=1.645z_{0.95}=1.645 and z0.9=1.282z_{0.9}=1.282 the bound is

    n≥((1.645+1.282)⋅20.5)2=11.7082=137.1,n\geq\Bigl{(}\frac{(1.645+1.282)\cdot 2}{0.5}\Bigr{)}^{2}=11.708^{2}=137.1,

    and thus n=138n=138 observations are needed.

  3. 3.
    ​

    The bound is proportional to σ2/(θ1−θ0)2\sigma^{2}/(\theta_{1}-\theta_{0})^{2}. Halving the difference θ1−θ0\theta_{1}-\theta_{0} multiplies the required sample size by four, to n=549n=549 in the numerical example (the bound becomes 548.3548.3), and doubling σ\sigma has the same effect. Detecting a difference half as large costs four times as many observations.

∎

Solution to exercise 20.2.
  1. 1.
    ​

    Under H0H_{0} the total TT is binomially distributed with parameters 1010 and 1/21/2. Thus, the test size is

    β⁢(1/2)=ℙ1/2⁢(T≥c)=1210⁢∑k=c10(10k).\beta(1/2)=\mathbb{P}_{1/2}(T\geq c)=\frac{1}{2^{10}}\sum_{k=c}^{10}\binom{10}% {k}.
  2. 2.
    ​

    We can compute the binomial coefficients (1010)=1\binom{10}{10}=1, (109)=10\binom{10}{9}=10 and (108)=45\binom{10}{8}=45 to find the sizes 1/1024=0.00101/1024=0.0010 for c=10c=10, 11/1024=0.010711/1024=0.0107 for c=9c=9 and 56/1024=0.054756/1024=0.0547 for c=8c=8. The smallest value cc with size less than or equal to 0.050.05 is c=9c=9 with size 0.01070.0107. Since TT can only take the values 0,…,100,\dots,10, the size can only take the eleven values computed in the first part of the question, and 0.050.05 is not one of these values (the size jumps from 0.01070.0107 to 0.05470.0547 between c=9c=9 and c=8c=8).

  3. 3.
    ​

    The power at pp is ℙp⁢(T≥9)=10⁢p9⁢(1−p)+p10=p9⁢(10−9⁢p)\mathbb{P}_{p}(T\geq 9)=10p^{9}(1-p)+p^{10}=p^{9}(10-9p). This gives the values 0.1490.149 for p=0.7p=0.7, 0.3760.376 for p=0.8p=0.8 and 0.7360.736 for p=0.9p=0.9. A coin which shows heads four times in five experiments is detected in fewer than 40%40\% of the experiments; for ten experiments the test can only detect large deviations from fairness with reasonable probability, and a failure to reject the hypothesis is not very meaningful.

  4. 4.
    ​

    The pp-value is the probability of the event that the total under H0H_{0} is larger than or equal to the observed value, i.e. the probability of ℙ1/2⁢(T≥9)=11/1024=0.0107\mathbb{P}_{1/2}(T\geq 9)=11/1024=0.0107. Since this probability is less than or equal to 0.050.05, we can reject H0H_{0} at a significance level of 0.050.05. This agrees with the critical region T≥9T\geq 9 from the second part of the question.

∎

Solution to exercise 20.3.
  1. 1.
    ​

    The maximum of the distribution is at most mm, if and only if all observations are at most mm. Since the observations are independent, for 0≤m≤θ0\leq m\leq\theta we have

    ℙθ⁢(M≤m)=∏i=1nℙθ⁢(Xi≤m)=(mθ)n,\mathbb{P}_{\theta}(M\leq m)=\prod_{i=1}^{n}\mathbb{P}_{\theta}(X_{i}\leq m)=% \Bigl{(}\frac{m}{\theta}\Bigr{)}^{n},

    This is also the integral of the density from example 8.10.

  2. 2.
    ​

    For c≤θ0c\leq\theta_{0} the size is ℙθ0⁢(M>c)=1−(c/θ0)n\mathbb{P}_{\theta_{0}}(M>c)=1-(c/\theta_{0})^{n}, as found in the first part. Setting this equal to α\alpha gives c=θ0⁢(1−α)1/nc=\theta_{0}(1-\alpha)^{1/n}.

  3. 3.
    ​

    For θ>θ0≥c\theta>\theta_{0}\geq c the power function is

    β⁢(θ)=ℙθ⁢(M>c)=1−(cθ)n=1−(1−α)⁢(θ0θ)n,\beta(\theta)=\mathbb{P}_{\theta}(M>c)=1-\Bigl{(}\frac{c}{\theta}\Bigr{)}^{n}=% 1-(1-\alpha)\Bigl{(}\frac{\theta_{0}}{\theta}\Bigr{)}^{n},

    This function increases from α\alpha for θ=θ0\theta=\theta_{0} to 11 as θ→∞\theta\to\infty. For n=5n=5, θ0=1\theta_{0}=1 and α=0.05\alpha=0.05 we have c=0.951/5=0.990c=0.95^{1/5}=0.990, and for θ=1.2\theta=1.2 the power is 1−0.95/1.25=1−0.95⋅0.402=0.6181-0.95/1.2^{5}=1-0.95\cdot 0.402=0.618.

  4. 4.
    ​

    Under H0H_{0} all observations are in the interval [0,θ0][0,\theta_{0}] and thus we have ℙθ0⁢(M>θ0)=0\mathbb{P}_{\theta_{0}}(M>\theta_{0})=0: The second test has size 0 and is a level α\alpha test for every α\alpha. The power function of this test is ℙθ⁢(M>θ0)=1−(θ0/θ)n\mathbb{P}_{\theta}(M>\theta_{0})=1-(\theta_{0}/\theta)^{n} for θ>θ0\theta>\theta_{0}. This function equals 0.5980.598 for θ=1.2\theta=1.2 in the numerical example. The difference between the two power functions is α⁢(θ0/θ)n\alpha(\theta_{0}/\theta)^{n}, and this is positive for every alternative, i.e. the first test is more powerful everywhere on H1H_{1}, but the first test has type I error probability α\alpha which the second test avoids. In the context of section 20.3, the first test is to be preferred: we have agreed on a level α\alpha, and a test which does not use the level would be a waste of power. The power gain is at most α\alpha, and decreases quickly as θ\theta increases, whereas for the second test a rejection is certain if the observation is above θ0\theta_{0}, which is impossible under H0H_{0}, and so we can reasonably assume that this is worth the small loss of power. The example shows that the choice of level is a matter for the statistician, and is not a mathematical necessity.

∎

Solution to exercise 20.4.
  1. 1.
    ​

    The observed value of the standardised statistic is z=16⁢(52.3−50)/4=2.3z=\sqrt{16}\,(52.3-50)/4=2.3 and the one-sided pp-value is p=1−Φ⁢(2.3)=0.0107p=1-\Phi(2.3)=0.0107. Since 0.0107≤0.050.0107\leq 0.05 we can reject H0H_{0} at level 0.050.05, but since 0.0107>0.010.0107>0.01 we cannot reject at level 0.010.01.

  2. 2.
    ​

    The two-sided pp-value is 2⁢(1−Φ⁢(2.3))=0.02142\bigl{(}1-\Phi(2.3)\bigr{)}=0.0214. Again we reject at level 0.050.05 and we cannot reject at level 0.010.01.

  3. 3.
    ​

    The mean θ\theta is a fixed, unknown number and not a random variable, so we cannot determine a probability for the statement θ=50\theta=50. The pp-value is a probability about the data, assuming that θ=50\theta=50. A correct interpretation is: if the true mean were 5050, then a sample mean of at least 52.352.3 would occur with probability 0.0110.011, and thus the data are significant at level 0.050.05 but not at level 0.010.01.

  4. 4.
    ​

    Under H0H_{0} the statistic ZZ is standard normally distributed. Since Φ\Phi is continuous and strictly increasing, for u∈(0,1)u\in(0,1) we have

    ℙθ0⁢(1−Φ⁢(Z)≤u)\displaystyle\mathbb{P}_{\theta_{0}}\bigl{(}1-\Phi(Z)\leq u\bigr{)}
    =ℙθ0⁢(Φ⁢(Z)≥1−u)\displaystyle=\mathbb{P}_{\theta_{0}}\bigl{(}\Phi(Z)\geq 1-u\bigr{)}
    =ℙθ0⁢(Z≥z1−u)=1−Φ⁢(z1−u)=u,\displaystyle=\mathbb{P}_{\theta_{0}}(Z\geq z_{1-u})=1-\Phi(z_{1-u})=u,

    and thus the pp-value is uniformly distributed on (0,1)(0,1) under H0H_{0}. In particular we have ℙθ0⁢(p⁢(X)≤α)=α\mathbb{P}_{\theta_{0}}\bigl{(}p(X)\leq\alpha\bigr{)}=\alpha, which is the statement that the test has size α\alpha.

∎

C.15 The Neyman–Pearson Lemma

Solution to exercise 22.1.
  1. 1.
    ​

    The joint probability weights are f⁢(x;θ)=θs⁢e−n⁢θ/∏ixi!f(x;\theta)=\theta^{s}e^{-n\theta}/\prod_{i}x_{i}! with s=∑ixis=\sum_{i}x_{i}. Thus the likelihood ratio is

    R⁢(x)=(θ1θ0)s⁢e−n⁢(θ1−θ0).R(x)=\Bigl{(}\frac{\theta_{1}}{\theta_{0}}\Bigr{)}^{s}e^{-n(\theta_{1}-\theta_% {0})}.

    Since θ1/θ0>1\theta_{1}/\theta_{0}>1, this function is strictly increasing as a function of ss and the inequality R⁢(x)>kR(x)>k is equivalent to s>cs>c with c=(log⁡k+n⁢(θ1−θ0))/log⁡(θ1/θ0)c=\bigl{(}\log k+n(\theta_{1}-\theta_{0})\bigr{)}/\log(\theta_{1}/\theta_{0}). By theorem 22.2 the most powerful test at this level is the one which rejects when S>cS>c. The size of this test, ℙθ0⁢(S>c)\mathbb{P}_{\theta_{0}}(S>c), can be found using the Poisson distribution with mean n⁢θ0n\theta_{0}, which contains the parameters nn and θ0\theta_{0} only. For a given size the threshold cc is the same for all θ1>θ0\theta_{1}>\theta_{0}. By proposition 22.4 the test is uniformly most powerful against H1:θ>θ0H_{1}\colon\theta>\theta_{0}.

  2. 2.
    ​

    Under H0H_{0}, the sum SS is Poisson-distributed with mean 55 and thus we have ℙθ0⁢(S>8)=1−0.9319=0.0681>0.05\mathbb{P}_{\theta_{0}}(S>8)=1-0.9319=0.0681>0.05, whereas ℙθ0⁢(S>9)=1−0.9682=0.0318≤0.05\mathbb{P}_{\theta_{0}}(S>9)=1-0.9682=0.0318\leq 0.05. The smallest acceptable threshold for the test is c=9c=9: we can reject S≥10S\geq 10 and the size of this test is 0.03180.0318.

  3. 3.
    ​

    Under θ1=2\theta_{1}=2, the sum is Poisson-distributed with mean 1010 and thus the power of the test is ℙθ1⁢(S>9)=1−0.4579=0.5421\mathbb{P}_{\theta_{1}}(S>9)=1-0.4579=0.5421.

  4. 4.
    ​

    The size of the test {S>c}\{S>c\} is 1−ℙθ0⁢(S≤c)1-\mathbb{P}_{\theta_{0}}(S\leq c). This value can only take the values listed for c=0,1,2,…c=0,1,2,\dots and the jumps between these values are from c=8c=8 to c=9c=9, i.e. 0.06810.0681 to 0.03180.0318. The value 0.050.05 is not one of the possible values. The test from part (b) has size 0.03180.0318 and the lemma shows that this is the most powerful test at level 0.03180.0318. The test is also a test of level 0.050.05, but the lemma does not show that this is the most powerful test among all tests with level 0.050.05. As the remark in section 22.2 points out, a randomised test could have size exactly 0.050.05 and higher power.

∎

Solution to exercise 22.2.
  1. 1.
    ​

    The joint probability weights are f⁢(x;p)=ps⁢(1−p)n−sf(x;p)=p^{s}(1-p)^{n-s} with s=∑ixis=\sum_{i}x_{i}. Under H0H_{0}, these weights equal 2−n2^{-n} for all xx. Thus we have

    R⁢(x)=p1s⁢(1−p1)n−s2−n=(p11−p1)s⁢(2⁢(1−p1))n,R(x)=\frac{p_{1}^{s}(1-p_{1})^{n-s}}{2^{-n}}=\Bigl{(}\frac{p_{1}}{1-p_{1}}% \Bigr{)}^{s}\bigl{(}2(1-p_{1})\bigr{)}^{n},

    as claimed. For p1>1/2p_{1}>1/2, we have p1/(1−p1)>1p_{1}/(1-p_{1})>1 and thus R⁢(x)R(x) is strictly increasing in ss. The inequality R⁢(x)>kR(x)>k is equivalent to s>cs>c for some cc. By theorem 22.2, the most powerful test is to reject when S>cS>c, where cc is determined by the distribution of SS under H0H_{0}, i.e. the binomial distribution with parameters nn and 1/21/2. Since this distribution does not depend on p1p_{1}, the same region of p1>1/2p_{1}>1/2 is optimal for all alternatives and by proposition 22.4 the test from exercise 20.2 is uniformly most powerful against H1:p>1/2H_{1}\colon p>1/2.

  2. 2.
    ​

    Since ℙ1/2⁢(S≥14)=0.0577>0.05\mathbb{P}_{1/2}(S\geq 14)=0.0577>0.05 and ℙ1/2⁢(S≥15)=0.0207≤0.05\mathbb{P}_{1/2}(S\geq 15)=0.0207\leq 0.05, the largest critical region which does not exceed size 0.050.05 is {S≥15}\{S\geq 15\}. This region has size 0.02070.0207. The power of this test at p1=0.7p_{1}=0.7 is ℙ0.7⁢(S≥15)=1−0.5836=0.4164\mathbb{P}_{0.7}(S\geq 15)=1-0.5836=0.4164. For n=10n=10, the corresponding region {S≥9}\{S\geq 9\} from exercise 20.2 has smaller size 11/1024=0.010711/1024=0.0107 but the power at p1=0.7p_{1}=0.7 is 10⋅0.79⋅0.3+0.710=0.149310\cdot 0.7^{9}\cdot 0.3+0.7^{10}=0.1493. By doubling the sample size, we can use a test with size closer to the nominal value 0.050.05 and nearly triple the power at this alternative.

  3. 3.
    ​

    For p1<1/2p_{1}<1/2, we have p1/(1−p1)<1p_{1}/(1-p_{1})<1 and thus R⁢(x)R(x) is strictly decreasing in ss. The inequality R⁢(x)>kR(x)>k is now equivalent to s<cs<c and the most powerful test will be a test which rejects for small values of SS. For n=20n=20, with the same size 0.02070.0207, the test will reject for S≤5S\leq 5, by the symmetry of the binomial distribution with parameter 1/21/2. Thus, the two one-sided problems will result in tests which reject on the opposite tails of the distribution of SS and there can be no single critical region which satisfies both tests.

∎

Solution to exercise 22.3.
  1. 1.
    ​

    For x1,…,xn≥0x_{1},\dots,x_{n}\geq 0 we have

    f⁢(x;θ0)=θ0−n⁢𝟏{m≤θ0}andf⁢(x;θ1)=θ1−n⁢𝟏{m≤θ1},f(x;\theta_{0})=\theta_{0}^{-n}\mathbf{1}_{\{m\leq\theta_{0}\}}\qquad\text{and% }\qquad f(x;\theta_{1})=\theta_{1}^{-n}\mathbf{1}_{\{m\leq\theta_{1}\}},

    and since θ0<θ1\theta_{0}<\theta_{1} there are three cases. If m≤θ0m\leq\theta_{0}, both densities are positive and

    f⁢(x;θ1)=θ1−n=(θ0/θ1)n⁢θ0−n=k⁢f⁢(x;θ0).f(x;\theta_{1})=\theta_{1}^{-n}=(\theta_{0}/\theta_{1})^{n}\theta_{0}^{-n}=k\,% f(x;\theta_{0}).

    If θ0<m≤θ1\theta_{0}<m\leq\theta_{1}, we have f⁢(x;θ0)=0<f⁢(x;θ1)f(x;\theta_{0})=0<f(x;\theta_{1}), and thus f⁢(x;θ1)>k⁢f⁢(x;θ0)f(x;\theta_{1})>k\,f(x;\theta_{0}). If m>θ1m>\theta_{1}, both densities vanish and equality holds. Consequently

    {x∣f⁢(x;θ1)>k⁢f⁢(x;θ0)}\displaystyle\{x\mid f(x;\theta_{1})>k\,f(x;\theta_{0})\}
    ={θ0<m≤θ1},\displaystyle=\{\theta_{0}<m\leq\theta_{1}\},
    {x∣f⁢(x;θ1)=k⁢f⁢(x;θ0)}\displaystyle\{x\mid f(x;\theta_{1})=k\,f(x;\theta_{0})\}
    ={m≤θ0}∪{m>θ1}.\displaystyle=\{m\leq\theta_{0}\}\cup\{m>\theta_{1}\}.
  2. 2.
    ​

    Since 0<c<θ00<c<\theta_{0}, the size of CC is ℙθ0⁢(M>c)=1−(c/θ0)n=1−(1−α)=α\mathbb{P}_{\theta_{0}}(M>c)=1-(c/\theta_{0})^{n}=1-(1-\alpha)=\alpha. The region C={m>c}C=\{m>c\} contains the set {θ0<m≤θ1}\{\theta_{0}<m\leq\theta_{1}\} of part (a), because c<θ0c<\theta_{0}, and it is contained in the union of the two sets of part (a), which is the whole sample space. Thus CC satisfies the condition of theorem 22.2 with k=(θ0/θ1)nk=(\theta_{0}/\theta_{1})^{n} and by the first statement of the theorem it is most powerful at level α\alpha. The power of this test is

    ℙθ1⁢(M>c)=1−(cθ1)n=1−(1−α)⁢(θ0θ1)n.\mathbb{P}_{\theta_{1}}(M>c)=1-\Bigl{(}\frac{c}{\theta_{1}}\Bigr{)}^{n}=1-(1-% \alpha)\Bigl{(}\frac{\theta_{0}}{\theta_{1}}\Bigr{)}^{n}.

    The threshold c=θ0⁢(1−α)1/nc=\theta_{0}(1-\alpha)^{1/n} does not depend on θ1\theta_{1} and the argument from above applies to all θ1>θ0\theta_{1}>\theta_{0} with the corresponding kk and thus, by proposition 22.4, the test is uniformly most powerful against H1:θ>θ0H_{1}\colon\theta>\theta_{0}.

  3. 3.
    ​

    The two parts of C′′C^{\prime\prime} are disjoint and thus the size of this region is ℙθ0⁢(M>θ0)+ℙθ0⁢(M≤θ0⁢α1/n)=0+α=α\mathbb{P}_{\theta_{0}}(M>\theta_{0})+\mathbb{P}_{\theta_{0}}(M\leq\theta_{0}% \alpha^{1/n})=0+\alpha=\alpha. The power of this test is

    ℙθ1⁢(M>θ0)+ℙθ1⁢(M≤θ0⁢α1/n)=1−(θ0θ1)n+α⁢(θ0θ1)n=1−(1−α)⁢(θ0θ1)n,\mathbb{P}_{\theta_{1}}(M>\theta_{0})+\mathbb{P}_{\theta_{1}}(M\leq\theta_{0}% \alpha^{1/n})=1-\Bigl{(}\frac{\theta_{0}}{\theta_{1}}\Bigr{)}^{n}+\alpha\Bigl{% (}\frac{\theta_{0}}{\theta_{1}}\Bigr{)}^{n}=1-(1-\alpha)\Bigl{(}\frac{\theta_{% 0}}{\theta_{1}}\Bigr{)}^{n},

    This is the same power as for CC. Thus both regions are most powerful at level α\alpha and, by the second statement of theorem 22.2, they can only differ on the boundary set. And indeed they do: the two regions differ on the sets C∖C′′={c<m≤θ0}C\setminus C^{\prime\prime}=\{c<m\leq\theta_{0}\} and C′′∖C={m≤θ0⁢α1/n}C^{\prime\prime}\setminus C=\{m\leq\theta_{0}\alpha^{1/n}\}, both of which are contained in {m≤θ0}\{m\leq\theta_{0}\} where f⁢(x;θ1)=k⁢f⁢(x;θ0)f(x;\theta_{1})=k\,f(x;\theta_{0}). The argument in the uniqueness comment in section 22.2 used the fact that the likelihood ratio has a continuous distribution, but in this case this is not true: on the set {M≤θ0}\{M\leq\theta_{0}\}, where an event with probability 1−α1-\alpha occurs under H0H_{0}, the ratio R⁢(X)R(X) is constant and thus the boundary set has positive probability, despite the model having densities.

∎

Solution to exercise 22.4.

Since TT is sufficient, we can use the factorisation theorem 8.2 to find functions gg and hh such that

f⁢(x;θ)=g⁢(T⁢(x);θ)⁢h⁢(x)f(x;\theta)=g\bigl{(}T(x);\theta\bigr{)}\,h(x)

for all xx and all parameter values θ\theta. Now let us choose an xx with f⁢(x;θ0)>0f(x;\theta_{0})>0. Then neither of the two factors in the expression for f⁢(x;θ0)=g⁢(T⁢(x);θ0)⁢h⁢(x)f(x;\theta_{0})=g\bigl{(}T(x);\theta_{0}\bigr{)}h(x) can be zero and we get

R⁢(x)=f⁢(x;θ1)f⁢(x;θ0)=g⁢(T⁢(x);θ1)⁢h⁢(x)g⁢(T⁢(x);θ0)⁢h⁢(x)=g⁢(T⁢(x);θ1)g⁢(T⁢(x);θ0).R(x)=\frac{f(x;\theta_{1})}{f(x;\theta_{0})}=\frac{g\bigl{(}T(x);\theta_{1}% \bigr{)}\,h(x)}{g\bigl{(}T(x);\theta_{0}\bigr{)}\,h(x)}=\frac{g\bigl{(}T(x);% \theta_{1}\bigr{)}}{g\bigl{(}T(x);\theta_{0}\bigr{)}}.

Since we can cancel the factor h⁢(x)h(x) on both sides, the right-hand side only depends on xx via T⁢(x)T(x). This completes the proof. ∎

C.16 Likelihood-Ratio Tests

Here we solve the exercises at the end of lecture 23.

Solution to exercise 23.1.
  1. 1.
    ​

    By example 5.4 the log-likelihood

    ℓ⁢(μ,σ2)=−n2⁢log⁡(2⁢π⁢σ2)−12⁢σ2⁢∑i=1n(Xi−μ)2\ell(\mu,\sigma^{2})=-\frac{n}{2}\log(2\pi\sigma^{2})-\frac{1}{2\sigma^{2}}% \sum_{i=1}^{n}(X_{i}-\mu)^{2}

    is maximised over Θ\Theta at μ^=X¯\hat{\mu}=\bar{X} and σ^2=Q/n\hat{\sigma}^{2}=Q/n. Substituting these values, the sum of squares becomes QQ and the second term becomes Q/(2⁢Q/n)=n/2Q/(2Q/n)=n/2, and thus the maximised log-likelihood is

    ℓ⁢(μ^,σ^2)=−n2⁢log⁡(2⁢π⁢Qn)−n2.\ell(\hat{\mu},\hat{\sigma}^{2})=-\frac{n}{2}\log\Bigl{(}\frac{2\pi Q}{n}\Bigr% {)}-\frac{n}{2}.
  2. 2.
    ​

    Under H0H_{0} the mean is fixed at μ0\mu_{0} and only v=σ2v=\sigma^{2} is free. Writing Q0=∑i(Xi−μ0)2Q_{0}=\sum_{i}(X_{i}-\mu_{0})^{2}, the function to maximise is −n2⁢log⁡(2⁢π⁢v)−Q0/(2⁢v)-\frac{n}{2}\log(2\pi v)-Q_{0}/(2v), which is the function of example 5.4 with Q0Q_{0} in place of QQ, and thus by the same derivative computation it is maximised at σ^02=Q0/n\hat{\sigma}_{0}^{2}=Q_{0}/n, with maximum −n2⁢log⁡(2⁢π⁢Q0/n)−n2-\frac{n}{2}\log(2\pi Q_{0}/n)-\frac{n}{2}. The identity (2.1) with μ0\mu_{0} in place of μ\mu gives Q0=Q+n⁢(X¯−μ0)2Q_{0}=Q+n(\bar{X}-\mu_{0})^{2}, and thus

    σ^02=Q+n⁢(X¯−μ0)2n.\hat{\sigma}_{0}^{2}=\frac{Q+n(\bar{X}-\mu_{0})^{2}}{n}.
  3. 3.
    ​

    By (23.1) and the two previous parts, the terms −n/2-n/2 cancel and we get

    W=2⁢(−n2⁢log⁡2⁢π⁢Qn+n2⁢log⁡2⁢π⁢Q0n)=n⁢log⁡Q0Q=n⁢log⁡(1+n⁢(X¯−μ0)2Q).W=2\Bigl{(}-\frac{n}{2}\log\frac{2\pi Q}{n}+\frac{n}{2}\log\frac{2\pi Q_{0}}{n% }\Bigr{)}=n\log\frac{Q_{0}}{Q}=n\log\Bigl{(}1+\frac{n(\bar{X}-\mu_{0})^{2}}{Q}% \Bigr{)}.

    Since Q=(n−1)⁢S2Q=(n-1)S^{2}, we have n⁢(X¯−μ0)2/Q=n⁢(X¯−μ0)2/((n−1)⁢S2)=t2/(n−1)n(\bar{X}-\mu_{0})^{2}/Q=n(\bar{X}-\mu_{0})^{2}/\bigl{(}(n-1)S^{2}\bigr{)}=t^{% 2}/(n-1), which gives the stated formula. The function t2↦n⁢log⁡(1+t2/(n−1))t^{2}\mapsto n\log\bigl{(}1+t^{2}/(n-1)\bigr{)} is strictly increasing, and thus the event W≥cW\geq c is the event |t|≥c′|t|\geq c^{\prime} for a suitable c′c^{\prime}: the generalised likelihood-ratio test rejects for large values of |t||t|.

  4. 4.
    ​

    The full parameter θ=(μ,σ2)\theta=(\mu,\sigma^{2}) has k=2k=2 coordinates, and H0H_{0} fixes one of them, μ=μ0\mu=\mu_{0}, leaving σ2\sigma^{2} free, so that r=1r=1. The normal model satisfies the regularity conditions, and thus theorem 23.3 gives W→dχ12W\xrightarrow{\ \mathrm{d}\ }\chi^{2}_{1} under H0H_{0}. The approximate test rejects when W≥3.841W\geq 3.841, i.e. when t2≥(n−1)⁢(e3.841/n−1)t^{2}\geq(n-1)\bigl{(}e^{3.841/n}-1\bigr{)}. For n=10n=10 this is t2≥9⁢(e0.3841−1)=9⋅0.4683=4.215t^{2}\geq 9\,(e^{0.3841}-1)=9\cdot 0.4683=4.215, or |t|≥2.053|t|\geq 2.053, whereas the exact test rejects when |t|≥2.262|t|\geq 2.262. The approximate test rejects more often than the exact one, and its true size is ℙ⁢(|t9|≥2.053)=0.070\mathbb{P}(|t_{9}|\geq 2.053)=0.070 instead of the nominal 0.050.05: the chi-squared approximation is optimistic for small nn, because the t9t_{9} distribution has heavier tails than the normal limit which the approximation assumes. For large nn the two tests merge, since n⁢log⁡(1+t2/(n−1))→t2n\log\bigl{(}1+t^{2}/(n-1)\bigr{)}\to t^{2} by log⁡(1+x)≈x\log(1+x)\approx x and the tn−1t_{n-1} distribution tends to the standard normal one; for n=100n=100 the two thresholds are 1.9691.969 and 1.9841.984. Whenever the exact distribution is available, as here, it is the one to use, and lecture 29 does so.

∎

Solution to exercise 23.2.
  1. 1.
    ​

    Taking logarithms of the probability weights gives

    ℓ⁢(p)=log⁡n!n1!⁢⋯⁢nm!+∑j=1mnj⁢log⁡pj,\ell(p)=\log\frac{n!}{n_{1}!\cdots n_{m}!}+\sum_{j=1}^{m}n_{j}\log p_{j},

    where the first term does not depend on pp. To maximise ∑jnj⁢log⁡pj\sum_{j}n_{j}\log p_{j} subject to ∑jpj=1\sum_{j}p_{j}=1 we consider the Lagrangian ∑jnj⁢log⁡pj−λ⁢(∑jpj−1)\sum_{j}n_{j}\log p_{j}-\lambda\bigl{(}\sum_{j}p_{j}-1\bigr{)}. Setting its partial derivative with respect to pjp_{j} to zero gives nj/pj=λn_{j}/p_{j}=\lambda, i.e. pj=nj/λp_{j}=n_{j}/\lambda for every jj, and summing over jj and using the constraint we find 1=n/λ1=n/\lambda, so that λ=n\lambda=n and p^j=nj/n\hat{p}_{j}=n_{j}/n. This stationary point is the maximum, because ∑jnj⁢log⁡pj\sum_{j}n_{j}\log p_{j} is a concave function of pp and the constraint set is convex. If some njn_{j} is zero, the corresponding p^j=0\hat{p}_{j}=0 lies on the boundary of the parameter space; the formula still gives the maximum, since any probability assigned to an unobserved outcome only lowers the likelihood.

  2. 2.
    ​

    Since H0H_{0} is simple, the restricted maximiser is p0p^{0}, and by (23.1) we get

    W=2⁢(ℓ⁢(p^)−ℓ⁢(p0))=2⁢(∑j=1mnj⁢log⁡njn−∑j=1mnj⁢log⁡pj0)=2⁢∑j=1mnj⁢log⁡njn⁢pj0,W=2\bigl{(}\ell(\hat{p})-\ell(p^{0})\bigr{)}=2\Bigl{(}\sum_{j=1}^{m}n_{j}\log% \frac{n_{j}}{n}-\sum_{j=1}^{m}n_{j}\log p_{j}^{0}\Bigr{)}=2\sum_{j=1}^{m}n_{j}% \log\frac{n_{j}}{n\,p_{j}^{0}},

    where the combinatorial constant cancels. A term with nj=0n_{j}=0 contributes nothing to either sum, which is the convention 0⁢log⁡0=00\log 0=0.

  3. 3.
    ​

    The parameter space Θ={p∣pj>0,∑jpj=1}\Theta=\{p\mid p_{j}>0,\ \sum_{j}p_{j}=1\} is described by the m−1m-1 free coordinates p1,…,pm−1p_{1},\dots,p_{m-1}, with pmp_{m} determined by the constraint, and thus has dimension k=m−1k=m-1. The null hypothesis fixes all of them, so that r=k=m−1r=k=m-1. The counts are the sums of nn i.i.d. indicator vectors, one per repetition, whose distribution is a full-rank exponential family with support not depending on pp, so that the regularity conditions hold at every interior point p0p^{0}. Theorem 23.3 therefore gives W→dχm−12W\xrightarrow{\ \mathrm{d}\ }\chi^{2}_{m-1} under H0H_{0}.

  4. 4.
    ​

    Under H0:pj0=1/6H_{0}\colon p_{j}^{0}=1/6 the expected counts are n⁢pj0=10np_{j}^{0}=10 for every face. The terms nj⁢log⁡(nj/10)n_{j}\log(n_{j}/10) are

    13⁢log⁡1.3=3.411,8⁢log⁡0.8=−1.785,11⁢log⁡1.1=1.048,\displaystyle 13\log 1.3=3.411,\quad 8\log 0.8=-1.785,\quad 11\log 1.1=1.048,
    6⁢log⁡0.6=−3.065,12⁢log⁡1.2=2.188,10⁢log⁡1=0,\displaystyle 6\log 0.6=-3.065,\quad 12\log 1.2=2.188,\quad 10\log 1=0,

    with sum 1.7971.797, and thus w=2⋅1.797=3.59w=2\cdot 1.797=3.59. Since 3.59<11.07=χ5,0.9523.59<11.07=\chi^{2}_{5,0.95}, we do not reject H0H_{0} at level 0.050.05; the approximate pp-value is ℙ⁢(χ52≥3.59)=0.61\mathbb{P}(\chi^{2}_{5}\geq 3.59)=0.61. The observed counts are entirely consistent with a fair die.

  5. 5.
    ​

    Write xj=dj/(n⁢pj0)x_{j}=d_{j}/(np_{j}^{0}), so that nj=n⁢pj0⁢(1+xj)n_{j}=np_{j}^{0}(1+x_{j}) and nj/(n⁢pj0)=1+xjn_{j}/(np_{j}^{0})=1+x_{j}. Using log⁡(1+x)=x−x2/2+O⁢(x3)\log(1+x)=x-x^{2}/2+O(x^{3}) we get

    nj⁢log⁡njn⁢pj0\displaystyle n_{j}\log\frac{n_{j}}{np_{j}^{0}}
    =n⁢pj0⁢(1+xj)⁢(xj−xj22+O⁢(xj3))\displaystyle=np_{j}^{0}\,(1+x_{j})\Bigl{(}x_{j}-\frac{x_{j}^{2}}{2}+O(x_{j}^{% 3})\Bigr{)}
    =n⁢pj0⁢(xj+xj22+O⁢(xj3))=dj+dj22⁢n⁢pj0+O⁢(xj3),\displaystyle=np_{j}^{0}\Bigl{(}x_{j}+\frac{x_{j}^{2}}{2}+O(x_{j}^{3})\Bigr{)}% =d_{j}+\frac{d_{j}^{2}}{2np_{j}^{0}}+O(x_{j}^{3}),

    since n⁢pj0⁢xj=djnp_{j}^{0}x_{j}=d_{j} and n⁢pj0⁢xj2=dj2/(n⁢pj0)np_{j}^{0}x_{j}^{2}=d_{j}^{2}/(np_{j}^{0}). Summing over jj, the linear terms add up to ∑jdj=0\sum_{j}d_{j}=0, and doubling the result gives

    W≈∑j=1mdj2n⁢pj0=∑j=1m(nj−n⁢pj0)2n⁢pj0,W\approx\sum_{j=1}^{m}\frac{d_{j}^{2}}{np_{j}^{0}}=\sum_{j=1}^{m}\frac{(n_{j}-% np_{j}^{0})^{2}}{np_{j}^{0}},

    up to terms of third order in the relative deviations xjx_{j}. For the die the deviations from 1010 are 3,−2,1,−4,2,03,-2,1,-4,2,0, and Pearson’s statistic is (9+4+1+16+4+0)/10=3.4(9+4+1+16+4+0)/10=3.4, close to the value 3.593.59 of WW; the two statistics agree to the extent that the relative deviations are small, which for these data they are.

∎

Solution to exercise 23.3.
  1. 1.
    ​

    By example 4.2 the likelihood is L⁢(p)=pS⁢(1−p)n−SL(p)=p^{S}(1-p)^{n-S}, so that ℓ⁢(p)=S⁢log⁡p+(n−S)⁢log⁡(1−p)\ell(p)=S\log p+(n-S)\log(1-p), and by example 5.3 it is maximised over (0,1)(0,1) at p^=S/n\hat{p}=S/n when 0<S<n0<S<n. Since H0H_{0} is simple, the restricted maximiser is p0p_{0}, and (23.1) gives

    W=2⁢(ℓ⁢(p^)−ℓ⁢(p0))=2⁢(S⁢log⁡p^p0+(n−S)⁢log⁡1−p^1−p0).W=2\bigl{(}\ell(\hat{p})-\ell(p_{0})\bigr{)}=2\Bigl{(}S\log\frac{\hat{p}}{p_{0% }}+(n-S)\log\frac{1-\hat{p}}{1-p_{0}}\Bigr{)}.

    For S=0S=0 or S=nS=n the supremum of LL over (0,1)(0,1) is approached at the boundary, where one of the two terms is absent, and the convention 0⁢log⁡0=00\log 0=0 gives the correct value.

  2. 2.
    ​

    The parameter is one-dimensional and H0H_{0} fixes it, so that k=r=1k=r=1; the Bernoulli model satisfies the regularity conditions for p0∈(0,1)p_{0}\in(0,1), as noted in lecture 16. Theorem 23.3 gives W→dχ12W\xrightarrow{\ \mathrm{d}\ }\chi^{2}_{1} under H0H_{0}, and the approximate level α\alpha test rejects when W≥χ1,1−α2W\geq\chi^{2}_{1,1-\alpha}, which for α=0.05\alpha=0.05 means W≥3.841W\geq 3.841.

  3. 3.
    ​

    We have p^=32/50=0.64\hat{p}=32/50=0.64 and p0=0.5p_{0}=0.5, and thus

    w=2⁢(32⁢log⁡1.28+18⁢log⁡0.72)=2⁢(7.900−5.913)=3.97.w=2\bigl{(}32\log 1.28+18\log 0.72\bigr{)}=2\,(7.900-5.913)=3.97.

    Since 3.97≥3.8413.97\geq 3.841, we reject H0:p=1/2H_{0}\colon p=1/2 at level 0.050.05. The approximate pp-value is ℙ⁢(χ12≥3.97)=2⁢(1−Φ⁢(3.97))=2⁢(1−Φ⁢(1.99))=0.047\mathbb{P}(\chi^{2}_{1}\geq 3.97)=2\bigl{(}1-\Phi(\sqrt{3.97})\bigr{)}=2\bigl{% (}1-\Phi(1.99)\bigr{)}=0.047, just below 0.050.05, so that the evidence against a fair coin is real but not strong.

  4. 4.
    ​

    The map g⁢(p)=log⁡(p/(1−p))g(p)=\log\bigl{(}p/(1-p)\bigr{)} is a strictly increasing bijection from (0,1)(0,1) onto ℝ\mathbb{R} with inverse p=eψ/(1+eψ)p=e^{\psi}/(1+e^{\psi}), and the likelihood in the new parametrisation is Lψ⁢(ψ)=L⁢(g−1⁢(ψ))L_{\psi}(\psi)=L\bigl{(}g^{-1}(\psi)\bigr{)}. The null hypothesis ψ=ψ0\psi=\psi_{0} is the same set of distributions as p=p0p=p_{0} with p0=g−1⁢(ψ0)p_{0}=g^{-1}(\psi_{0}), and thus the numerator of the ratio is Lψ⁢(ψ0)=L⁢(p0)L_{\psi}(\psi_{0})=L(p_{0}) in both cases. Since the supremum supp∈(0,1)L⁢(p)\sup_{p\in(0,1)}L(p) is attained at the MLE p^\hat{p} by the first part, and gg is a bijection from (0,1)(0,1) onto ℝ\mathbb{R}, the two suprema supψ∈ℝLψ⁢(ψ)\sup_{\psi\in\mathbb{R}}L_{\psi}(\psi) and supp∈(0,1)L⁢(p)\sup_{p\in(0,1)}L(p) run over the same set of values, and thus the denominator is supψ∈ℝLψ⁢(ψ)=supp∈(0,1)L⁢(p)=L⁢(p^)\sup_{\psi\in\mathbb{R}}L_{\psi}(\psi)=\sup_{p\in(0,1)}L(p)=L(\hat{p}). Both ratios are therefore equal, and so are the two statistics WW. This is the invariance noted in section 23.1; by contrast the standard errors of exercise 17.1 do depend on the parametrisation.

∎

C.17 Confidence Intervals

Solution to exercise 25.1.
  1. 1.
    ​

    The function Q=2⁢θ⁢TQ=2\theta T is a function of the data and the parameter, and from exercise 13.1 we know that under ℙθ\mathbb{P}_{\theta} the distribution of QQ is χ2⁢n2\chi^{2}_{2n} for all θ\theta, independent of θ\theta. Thus, QQ is a pivot. Using a=χ2⁢n,α/22a=\chi^{2}_{2n,\alpha/2} and b=χ2⁢n,1−α/22b=\chi^{2}_{2n,1-\alpha/2} we find ℙθ⁢(a≤2⁢θ⁢T≤b)=1−α\mathbb{P}_{\theta}(a\leq 2\theta T\leq b)=1-\alpha and since the function QQ is increasing in θ\theta, the two inequalities can be solved for θ\theta directly:

    a≤2⁢θ⁢T≤b⟺χ2⁢n,α/222⁢T≤θ≤χ2⁢n,1−α/222⁢T.a\leq 2\theta T\leq b\quad\Longleftrightarrow\quad\frac{\chi^{2}_{2n,\alpha/2}% }{2T}\leq\theta\leq\frac{\chi^{2}_{2n,1-\alpha/2}}{2T}.

    Thus, the interval

    [χ2⁢n,α/22/(2⁢T),χ2⁢n,1−α/22/(2⁢T)]\bigl{[}\chi^{2}_{2n,\alpha/2}/(2T),\ \chi^{2}_{2n,1-\alpha/2}/(2T)\bigr{]}

    is a confidence interval for θ\theta with level 1−α1-\alpha and the probability of this interval covering the correct value is exactly 1−α1-\alpha.

  2. 2.
    ​

    We have 2⁢T=502T=50 and 2⁢n=202n=20 and thus the 95%95\% interval is

    [9.591/50, 34.170/50]=[0.192, 0.683].[9.591/50,\ 34.170/50]=[0.192,\ 0.683].
  3. 3.
    ​

    The MLE for the rate is θ^=1/x¯=10/25=0.4\hat{\theta}=1/\bar{x}=10/25=0.4 and the standard error of the MLE is 0.4/10=0.1260.4/\sqrt{10}=0.126. Thus the Wald interval is 0.4±1.96⋅0.1260.4\pm 1.96\cdot 0.126, which is the interval [0.152,0.648][0.152,0.648]. The exact interval is asymmetric about the estimate, extending further to the right than to the left. This is caused by the right skew of the distribution of θ^\hat{\theta} for small exponential samples; the Wald interval, being symmetric, extends to lower values than the exact interval does. At n=10n=10 the normal approximation used in the Wald interval is rough, as noted in section 17.5, and the exact interval should be used instead.

∎

Solution to exercise 25.2.
  1. 1.
    ​

    From example 8.10 we know that the maximum has density fM⁢(t)=n⁢tn−1/θnf_{M}(t)=nt^{n-1}/\theta^{n} for 0≤t≤θ0\leq t\leq\theta. Thus, for 0≤u≤10\leq u\leq 1 we have

    ℙθ⁢(Mθ≤u)=ℙθ⁢(M≤u⁢θ)=∫0u⁢θn⁢tn−1θn⁢dt=(u⁢θ)nθn=un.\mathbb{P}_{\theta}\Bigl{(}\frac{M}{\theta}\leq u\Bigr{)}=\mathbb{P}_{\theta}(% M\leq u\theta)=\int_{0}^{u\theta}\frac{nt^{n-1}}{\theta^{n}}\,\mathrm{d}t=% \frac{(u\theta)^{n}}{\theta^{n}}=u^{n}.

    The distribution function of M/θM/\theta does not depend on θ\theta and thus M/θM/\theta is a pivot.

  2. 2.
    ​

    Since M≤θM\leq\theta always holds, we have

    ℙθ⁢(M≤θ≤Mα1/n)=ℙθ⁢(α1/n≤Mθ≤1)=1−(α1/n)n=1−α,\mathbb{P}_{\theta}\Bigl{(}M\leq\theta\leq\frac{M}{\alpha^{1/n}}\Bigr{)}=% \mathbb{P}_{\theta}\Bigl{(}\alpha^{1/n}\leq\frac{M}{\theta}\leq 1\Bigr{)}=1-% \bigl{(}\alpha^{1/n}\bigr{)}^{n}=1-\alpha,

    where we used the first part of the solution. Thus, [M,M/α1/n][M,M/\alpha^{1/n}] is a confidence interval with level 1−α1-\alpha. For n=5n=5 and α=0.05\alpha=0.05 we have α1/n=0.050.2=0.549\alpha^{1/n}=0.05^{0.2}=0.549 and thus the interval is [3.2, 3.2/0.549]=[3.2,5.83][3.2,\ 3.2/0.549]=[3.2,5.83].

  3. 3.
    ​

    The condition for coverage is bn−an=1−αb^{n}-a^{n}=1-\alpha. This condition can be used to find a=(bn−1+α)1/na=(b^{n}-1+\alpha)^{1/n} as a function of bb and thus, by choosing the interval [M/b,M/a][M/b,M/a], we can find the length M⁢(1/a−1/b)M(1/a-1/b) of the interval. Taking derivatives we find bn−1=an−1⁢d⁢a/d⁢bb^{n-1}=a^{n-1}\,\mathrm{d}a/\mathrm{d}b and thus

    dd⁢b⁢(1a−1b)=−1a2⁢d⁢ad⁢b+1b2=−bn−1an+1+1b2=1b2⁢(1−(ba)n+1)<0,\frac{\mathrm{d}}{\mathrm{d}b}\Bigl{(}\frac{1}{a}-\frac{1}{b}\Bigr{)}=-\frac{1% }{a^{2}}\,\frac{\mathrm{d}a}{\mathrm{d}b}+\frac{1}{b^{2}}=-\frac{b^{n-1}}{a^{n% +1}}+\frac{1}{b^{2}}=\frac{1}{b^{2}}\Bigl{(}1-\Bigl{(}\frac{b}{a}\Bigr{)}^{n+1% }\Bigr{)}<0,

    since b>ab>a. The length of the interval decreases as bb increases and thus is smallest for the largest possible value b=1b=1, where a=α1/na=\alpha^{1/n}. This is the interval from the previous part of the solution. The reason for this result is that the density n⁢un−1nu^{n-1} of the pivot is increasing on [0,1][0,1] and thus the probability 1−α1-\alpha is concentrated on the shortest possible interval which is placed at the upper end of the density.

∎

Solution to exercise 25.3.
  1. 1.
    ​

    The standard error is se(θ^n)=1/n⁢ℐ︀⁢(X¯)=X¯/n\mathop{\mathrm{se}}\nolimits(\hat{\theta}_{n})=1/\sqrt{n\mathcal{I}(\bar{X})}% =\sqrt{\bar{X}/n}, and the Wald interval is X¯±z1−α/2⁢X¯/n\bar{X}\pm z_{1-\alpha/2}\sqrt{\bar{X}/n}. For the data we have x¯=1.5\bar{x}=1.5 and 1.5/20=0.274\sqrt{1.5/20}=0.274, and the 95%95\% interval is 1.5±1.96⋅0.2741.5\pm 1.96\cdot 0.274, which is [0.96,2.04][0.96,2.04].

  2. 2.
    ​

    In exercise 17.2 we found n⁢(p^0−p0)→dN⁢(0,e−2⁢θ⁢θ)\sqrt{n}\,(\hat{p}_{0}-p_{0})\xrightarrow{\ \mathrm{d}\ }N(0,e^{-2\theta}\theta), and replacing θ\theta by X¯\bar{X} in the variance gives the standard error se(p^0)=e−X¯⁢X¯/n\mathop{\mathrm{se}}\nolimits(\hat{p}_{0})=e^{-\bar{X}}\sqrt{\bar{X}/n}, in agreement with the delta-method formula |g′⁢(θ^n)|⁢se(θ^n)|g^{\prime}(\hat{\theta}_{n})|\,\mathop{\mathrm{se}}\nolimits(\hat{\theta}_{n}) for g⁢(θ)=e−θg(\theta)=e^{-\theta}. The approximate interval is e−X¯±z1−α/2⁢e−X¯⁢X¯/ne^{-\bar{X}}\pm z_{1-\alpha/2}\,e^{-\bar{X}}\sqrt{\bar{X}/n}. For the data we have p^0=e−1.5=0.223\hat{p}_{0}=e^{-1.5}=0.223 and se(p^0)=0.223⋅0.274=0.061\mathop{\mathrm{se}}\nolimits(\hat{p}_{0})=0.223\cdot 0.274=0.061, and the 95%95\% interval is 0.223±1.96⋅0.0610.223\pm 1.96\cdot 0.061, which is [0.103,0.343][0.103,0.343].

  3. 3.
    ​

    Write [L,U][L,U] for the interval of the first part. Since gg is decreasing, the event L≤θ≤UL\leq\theta\leq U is the same as the event g⁢(U)≤g⁢(θ)≤g⁢(L)g(U)\leq g(\theta)\leq g(L), and thus the interval [e−U,e−L][e^{-U},e^{-L}] covers p0p_{0} exactly when [L,U][L,U] covers θ\theta. Its coverage probability is that of the Wald interval, which converges to 1−α1-\alpha by proposition 25.7. For the data we get [e−2.04,e−0.96]=[0.130,0.383][e^{-2.04},e^{-0.96}]=[0.130,0.383]. The two intervals differ because gg is not linear: the delta-method interval is symmetric about p^0\hat{p}_{0}, while the transformed interval is shifted towards larger values and is asymmetric about p^0=0.223\hat{p}_{0}=0.223. Both have asymptotic coverage 1−α1-\alpha, and for large nn the two nearly coincide, since gg is close to linear over a short interval. A practical advantage of the transformed interval is that it always lies inside (0,1)(0,1), whereas the delta-method interval can extend below zero when p^0\hat{p}_{0} is small.

∎

Solution to exercise 25.4.
  1. 1.
    ​

    Under H0H_{0}, the test statistic n⁢(X¯−μ0)/S\sqrt{n}\,(\bar{X}-\mu_{0})/S is the pivot TT from example 25.4 and has tn−1t_{n-1} distribution, independent of the value of σ2\sigma^{2}. Since the tt distribution is symmetric, we have

    ℙμ0,σ2⁢(|T|>tn−1,1−α/2)\displaystyle\mathbb{P}_{\mu_{0},\sigma^{2}}\bigl{(}|T|>t_{n-1,1-\alpha/2}% \bigr{)}
    =ℙ⁢(T>tn−1,1−α/2)+ℙ⁢(T<−tn−1,1−α/2)\displaystyle=\mathbb{P}(T>t_{n-1,1-\alpha/2})+\mathbb{P}(T<-t_{n-1,1-\alpha/2})
    =α2+α2=α,\displaystyle=\frac{\alpha}{2}+\frac{\alpha}{2}=\alpha,

    and the test has size α\alpha.

  2. 2.
    ​

    The acceptance region of the test for μ0\mu_{0} is A⁢(μ0)={x∣|x¯−μ0|≤tn−1,1−α/2⁢s/n}A(\mu_{0})=\{x\mid|\bar{x}-\mu_{0}|\leq t_{n-1,1-\alpha/2}\,s/\sqrt{n}\}. By theorem 25.6, the set of values not rejected,

    C⁢(X)\displaystyle C(X)
    ={μ0||X¯−μ0|≤tn−1,1−α/2⁢Sn}\displaystyle=\Bigl{\{}\mu_{0}\mathrel{\big{|}}|\bar{X}-\mu_{0}|\leq t_{n-1,1-% \alpha/2}\,\frac{S}{\sqrt{n}}\Bigr{\}}
    =[X¯−tn−1,1−α/2⁢Sn,X¯+tn−1,1−α/2⁢Sn],\displaystyle=\Bigl{[}\,\bar{X}-t_{n-1,1-\alpha/2}\,\frac{S}{\sqrt{n}},\ \bar{% X}+t_{n-1,1-\alpha/2}\,\frac{S}{\sqrt{n}}\,\Bigr{]},

    is a confidence set with level 1−α1-\alpha and it is the tt-interval from example 25.4.

  3. 3.
    ​

    The 95%95\% tt-interval for the given data is [9.18,11.42][9.18,11.42]. By duality, H0:μ=μ0H_{0}\colon\mu=\mu_{0} is rejected at level 5%5\% if and only if μ0\mu_{0} is outside this interval. Thus, H0:μ=9H_{0}\colon\mu=9 is rejected since 9<9.189<9.18, and H0:μ=10H_{0}\colon\mu=10 is not rejected since 1010 is inside the interval.

  4. 4.
    ​

    The one-sided test does not reject μ0\mu_{0} if and only if n⁢(X¯−μ0)/S≤tn−1,1−α\sqrt{n}\,(\bar{X}-\mu_{0})/S\leq t_{n-1,1-\alpha}, i.e. if μ0≥X¯−tn−1,1−α⁢S/n\mu_{0}\geq\bar{X}-t_{n-1,1-\alpha}\,S/\sqrt{n}. Thus, the set of values not rejected is the one-sided interval [X¯−tn−1,1−α⁢S/n,∞)\bigl{[}\bar{X}-t_{n-1,1-\alpha}\,S/\sqrt{n},\ \infty\bigr{)}, whose left-hand boundary is a lower confidence bound for μ\mu with level 1−α1-\alpha. Under H0H_{0}, the test statistic is tn−1t_{n-1} distributed, and thus the test has size α\alpha and the bound has coverage probability exactly 1−α1-\alpha, by theorem 25.6.

∎

C.18 Bayesian Inference: Priors and Posteriors

Solution to exercise 26.1.
  1. 1.
    ​

    The prior mean is a/(a+b)=0.3a/(a+b)=0.3 and since a+b=10a+b=10 we have a=3a=3 and b=7b=7. From table A.2 we find that the variance of the Beta⁢(3,7)\text{Beta}(3,7) distribution is

    a⁢b(a+b)2⁢(a+b+1)=21100⋅11≈0.0191,\frac{ab}{(a+b)^{2}(a+b+1)}=\frac{21}{100\cdot 11}\approx 0.0191,

    and thus the prior standard deviation is approximately 0.1380.138.

  2. 2.
    ​

    Let μ=a/(a+b)\mu=a/(a+b) denote the prior mean and s=a+bs=a+b the equivalent sample size. Then we have a=μ⁢sa=\mu s and b=(1−μ)⁢sb=(1-\mu)s. Using the formula for the variance from the table, we get

    a⁢b(a+b)2⁢(a+b+1)=μ⁢(1−μ)s+1.\frac{ab}{(a+b)^{2}(a+b+1)}=\frac{\mu(1-\mu)}{s+1}.

    Given μ=0.3\mu=0.3 and a variance of 0.12=0.010.1^{2}=0.01 we find that we need s+1=0.21/0.01=21s+1=0.21/0.01=21, i.e. s=20s=20 and thus we have a=6a=6 and b=14b=14. The equivalent sample size is 2020: the smaller spread corresponds to a prior which gives twice the number of observations.

  3. 3.
    ​

    From example 26.5 we know that the posterior distribution is Beta⁢(3+9, 7+11)=Beta⁢(12,18)\text{Beta}(3+9,\,7+11)=\text{Beta}(12,18). The mean of this distribution is 12/30=0.412/30=0.4. In the form (26.3) we can write this as

    1030⋅0.3+2030⋅0.45=0.1+0.3=0.4,\frac{10}{30}\cdot 0.3+\frac{20}{30}\cdot 0.45=0.1+0.3=0.4,

    where x¯=9/20=0.45\bar{x}=9/20=0.45: The data have twice the weight as the prior, and the posterior mean is two thirds of the way between the prior mean and the sample proportion.

∎

Solution to exercise 26.2.
  1. 1.
    ​

    On the interval (0,1/2](0,1/2] we have 1−θ≥1/21-\theta\geq 1/2 and thus θ−1⁢(1−θ)−1≥θ−1\theta^{-1}(1-\theta)^{-1}\geq\theta^{-1} there. Since ∫01/2θ−1⁢dθ=∞\int_{0}^{1/2}\theta^{-1}\,\mathrm{d}\theta=\infty, the integral of the prior over (0,1)(0,1) is infinite and the prior is improper.

  2. 2.
    ​

    Consider the kernel θα−1⁢(1−θ)β−1\theta^{\alpha-1}(1-\theta)^{\beta-1} for real α\alpha and β\beta. Near θ=0\theta=0 the factor (1−θ)β−1(1-\theta)^{\beta-1} is bounded between two positive constants and thus the kernel is integrable near 0 if and only if θα−1\theta^{\alpha-1} is, which is equivalent to α>0\alpha>0. By symmetry the kernel is integrable near θ=1\theta=1 if and only if β>0\beta>0 and thus, since it is bounded away from the endpoints, the integral over (0,1)(0,1) is finite if and only if α>0\alpha>0 and β>0\beta>0, in which case it equals B⁢(α,β)B(\alpha,\beta). For the formal posterior we have α=T\alpha=T and β=n−T\beta=n-T and thus the posterior is proper if and only if T>0T>0 and n−T>0n-T>0, i.e. if and only if 0<T<n0<T<n. In this case it is the Beta⁢(T,n−T)\text{Beta}(T,n-T) distribution.

  3. 3.
    ​

    From table A.2 we find that the posterior mean is T/(T+n−T)=T/n=x¯T/(T+n-T)=T/n=\bar{x}. This is the maximum likelihood estimate from example 5.3. The Haldane prior is the choice of prior where the Bayes estimator and the maximum likelihood estimator coincide. This happens for an equivalent sample size of zero in (26.3).

  4. 4.
    ​

    For T=0T=0 the likelihood (1−θ)n(1-\theta)^{n} has its maximum at θ=0\theta=0 and does not vanish there, while the prior has a non-integrable singularity at θ=0\theta=0. Since the data does not contain any successes, the infinite mass of the prior near zero cannot be compensated and thus the integral of θ−1⁢(1−θ)n−1\theta^{-1}(1-\theta)^{n-1} is infinite: there is no posterior distribution in this case. In informal language, the improper prior suggests that θ\theta is very close to either 0 or 11 and a data set with no successes cannot exclude the first of these two possibilities.

∎

Solution to exercise 26.3.
  1. 1.
    ​

    We have

    f⁢(x|θ)\displaystyle f(x\mskip 1.0mu|\mskip 1.0mu\theta)
    =∏i=1n12⁢π⁢σ2⁢exp⁡(−(xi−θ)22⁢σ2)\displaystyle=\prod_{i=1}^{n}\frac{1}{\sqrt{2\pi\sigma^{2}}}\exp\Bigl{(}-\frac% {(x_{i}-\theta)^{2}}{2\sigma^{2}}\Bigr{)}
    =(2⁢π⁢σ2)−n/2⁢exp⁡(−12⁢σ2⁢∑i=1n(xi−θ)2).\displaystyle=(2\pi\sigma^{2})^{-n/2}\exp\Bigl{(}-\frac{1}{2\sigma^{2}}\sum_{i% =1}^{n}(x_{i}-\theta)^{2}\Bigr{)}.

    Using the identity (2.1) from lecture 2, which holds for all numbers in place of the mean, and using θ\theta, we find ∑i(xi−θ)2=∑i(xi−x¯)2+n⁢(x¯−θ)2\sum_{i}(x_{i}-\theta)^{2}=\sum_{i}(x_{i}-\bar{x})^{2}+n(\bar{x}-\theta)^{2}. Since the first term does not depend on θ\theta and since the factor (2⁢π⁢σ2)−n/2(2\pi\sigma^{2})^{-n/2} does not depend on θ\theta, we have

    f⁢(x|θ)∝exp⁡(−n⁢(θ−x¯)2/(2⁢σ2)).f(x\mskip 1.0mu|\mskip 1.0mu\theta)\propto\exp\bigl{(}-n(\theta-\bar{x})^{2}/(% 2\sigma^{2})\bigr{)}.
  2. 2.
    ​

    Since the prior density, as a function of θ\theta, is proportional to exp⁡(−(θ−m)2/(2⁢τ2))\exp\bigl{(}-(\theta-m)^{2}/(2\tau^{2})\bigr{)} and since we can group the terms in the exponent of the posterior density as powers of θ\theta, we get

    π⁢(θ|x)\displaystyle\pi(\theta\mskip 1.0mu|\mskip 1.0mux)
    ∝exp⁡(−n⁢(θ−x¯)22⁢σ2−(θ−m)22⁢τ2)\displaystyle\propto\exp\Bigl{(}-\frac{n(\theta-\bar{x})^{2}}{2\sigma^{2}}-% \frac{(\theta-m)^{2}}{2\tau^{2}}\Bigr{)}
    ∝exp⁡(−12⁢(nσ2+1τ2)⁢θ2+(n⁢x¯σ2+mτ2)⁢θ),\displaystyle\propto\exp\Bigl{(}-\frac{1}{2}\Bigl{(}\frac{n}{\sigma^{2}}+\frac% {1}{\tau^{2}}\Bigr{)}\theta^{2}+\Bigl{(}\frac{n\bar{x}}{\sigma^{2}}+\frac{m}{% \tau^{2}}\Bigr{)}\theta\Bigr{)},

    where we have omitted the terms which do not contain θ\theta. Using the abbreviations 1/τn2=1/τ2+n/σ21/\tau_{n}^{2}=1/\tau^{2}+n/\sigma^{2} and mn=τn2⁢(m/τ2+n⁢x¯/σ2)m_{n}=\tau_{n}^{2}\bigl{(}m/\tau^{2}+n\bar{x}/\sigma^{2}\bigr{)} we can write the exponent as −θ2/(2⁢τn2)+θ⁢mn/τn2-\theta^{2}/(2\tau_{n}^{2})+\theta m_{n}/\tau_{n}^{2} and completing the square we find

    −θ22⁢τn2+θ⁢mnτn2=−(θ−mn)22⁢τn2+mn22⁢τn2,-\frac{\theta^{2}}{2\tau_{n}^{2}}+\frac{\theta m_{n}}{\tau_{n}^{2}}=-\frac{(% \theta-m_{n})^{2}}{2\tau_{n}^{2}}+\frac{m_{n}^{2}}{2\tau_{n}^{2}},

    where we have omitted the last term which does not contain θ\theta. Thus we have π⁢(θ|x)∝exp⁡(−(θ−mn)2/(2⁢τn2))\pi(\theta\mskip 1.0mu|\mskip 1.0mux)\propto\exp\bigl{(}-(\theta-m_{n})^{2}/(2% \tau_{n}^{2})\bigr{)} and this is the kernel of the N⁢(mn,τn2)N(m_{n},\tau_{n}^{2}) density and, since the posterior is a density, this is the corresponding normal distribution. If we multiply out mnm_{n} we get the weighted form of the posterior given in proposition 26.6.

  3. 3.
    ​

    As τ→∞\tau\to\infty with nn fixed, the prior precision 1/τ21/\tau^{2} goes to zero and thus τn2→σ2/n\tau_{n}^{2}\to\sigma^{2}/n and mn→x¯m_{n}\to\bar{x}: the posterior converges to N⁢(x¯,σ2/n)N(\bar{x},\sigma^{2}/n), i.e. to the posterior under the flat prior discussed in section 26.4. As n→∞n\to\infty with τ\tau fixed, the weight (1/τ2)/(1/τ2+n/σ2)(1/\tau^{2})/(1/\tau^{2}+n/\sigma^{2}) of the prior mean goes to zero and thus mn−x¯→0m_{n}-\bar{x}\to 0 and τn2≤σ2/n\tau_{n}^{2}\leq\sigma^{2}/n goes to zero: the posterior is now concentrated around x¯\bar{x} and the prior is washed out.

∎

Solution to exercise 26.4.
  1. 1.
    ​

    Under ℙθ\mathbb{P}_{\theta}, the total number TT of successes is binomially distributed with 𝔼θ⁢(T)=n⁢θ\mathbb{E}_{\theta}(T)=n\theta and Varθ(T)=n⁢θ⁢(1−θ)\mathop{\mathrm{Var}}\nolimits_{\theta}(T)=n\theta(1-\theta). Since θ^B\hat{\theta}_{B} is an affine function of TT, we have

    𝔼θ⁢(θ^B)\displaystyle\mathbb{E}_{\theta}(\hat{\theta}_{B})
    =a+n⁢θa+b+nand\displaystyle=\frac{a+n\theta}{a+b+n}\qquad\text{and}
    bias(θ^B)\displaystyle\mathop{\mathrm{bias}}\nolimits(\hat{\theta}_{B})
    =a+n⁢θ−(a+b+n)⁢θa+b+n=a−(a+b)⁢θa+b+n,\displaystyle=\frac{a+n\theta-(a+b+n)\theta}{a+b+n}=\frac{a-(a+b)\theta}{a+b+n},

    which equals zero if and only if θ=a/(a+b)\theta=a/(a+b), the prior mean, and

    Varθ(θ^B)=n⁢θ⁢(1−θ)(a+b+n)2.\mathop{\mathrm{Var}}\nolimits_{\theta}(\hat{\theta}_{B})=\frac{n\theta(1-% \theta)}{(a+b+n)^{2}}.

    The bias is of order 1/n1/n and the variance is of order 1/n1/n, both tending to zero as n→∞n\to\infty. Thus, by corollary 2.11, the estimator is consistent, although for fixed nn it is biased for all θ≠a/(a+b)\theta\neq a/(a+b).

  2. 2.
    ​

    For a=b=1a=b=1 and n=10n=10 we have θ^B=(T+1)/12\hat{\theta}_{B}=(T+1)/12 and thus, by theorem 2.7,

    MSE(θ^B)=10⁢θ⁢(1−θ)+(1−2⁢θ)2144,MSE(X¯)=θ⁢(1−θ)10.\mathop{\mathrm{MSE}}\nolimits(\hat{\theta}_{B})=\frac{10\theta(1-\theta)+(1-2% \theta)^{2}}{144},\qquad\mathop{\mathrm{MSE}}\nolimits(\bar{X})=\frac{\theta(1% -\theta)}{10}.

    At θ=1/2\theta=1/2 the bias is zero and we have MSE(θ^B)=2.5/144≈0.0174\mathop{\mathrm{MSE}}\nolimits(\hat{\theta}_{B})=2.5/144\approx 0.0174 compared to MSE(X¯)=0.025\mathop{\mathrm{MSE}}\nolimits(\bar{X})=0.025. Thus, the Bayes estimator is better. At θ=0.1\theta=0.1 we have MSE(θ^B)=(0.9+0.64)/144≈0.0107\mathop{\mathrm{MSE}}\nolimits(\hat{\theta}_{B})=(0.9+0.64)/144\approx 0.0107 compared to MSE(X¯)=0.009\mathop{\mathrm{MSE}}\nolimits(\bar{X})=0.009 and thus X¯\bar{X} is better. Neither estimator has smaller mean squared error for all θ\theta: the Bayes estimator is better in the middle of the parameter range, where the prior mean is close to the truth, and the reduced variance does not compensate for the bias introduced by the shrinkage towards 1/21/2 at the ends of the range.

∎

Solution to exercise 26.5.

For the first part of the question we can take derivatives. The log-likelihood function for a single observation is ℓ0⁢(x−θ)\ell_{0}(x-\theta) and thus the score function is

∂∂θ⁢log⁡f⁢(x|θ)=∂∂θ⁢ℓ0⁢(x−θ)=−ℓ0′⁢(x−θ).\frac{\partial}{\partial\theta}\,\log f(x\mskip 1.0mu|\mskip 1.0mu\theta)=% \frac{\partial}{\partial\theta}\,\ell_{0}(x-\theta)=-\ell_{0}^{\prime}(x-% \theta).

Since Z=X−θZ=X-\theta, with density f0f_{0}, for all values of θ\theta, the Fisher information is

ℐ︀⁢(θ)=𝔼⁢(ℓ0′⁢(X−θ)2)=∫ℓ0′⁢(z)2⁢f0⁢(z)⁢dz.\mathcal{I}(\theta)=\mathbb{E}\bigl{(}\ell_{0}^{\prime}(X-\theta)^{2}\bigr{)}=% \int\ell_{0}^{\prime}(z)^{2}\,f_{0}(z)\,\mathrm{d}z.

This does not depend on θ\theta.

For the second part of the question we have log⁡f⁢(x|σ)=ℓ0⁢(x/σ)−log⁡σ\log f(x\mskip 1.0mu|\mskip 1.0mu\sigma)=\ell_{0}(x/\sigma)-\log\sigma and thus the score function is

∂∂σ⁢(ℓ0⁢(x/σ)−log⁡σ)=−ℓ0′⁢(x/σ)⁢xσ2−1σ=−1+z⁢ℓ0′⁢(z)σ\frac{\partial}{\partial\sigma}\bigl{(}\ell_{0}(x/\sigma)-\log\sigma\bigr{)}=-% \ell_{0}^{\prime}(x/\sigma)\,\frac{x}{\sigma^{2}}-\frac{1}{\sigma}=-\frac{1+z% \ell_{0}^{\prime}(z)}{\sigma}

where z=x/σz=x/\sigma. Since Z=X/σZ=X/\sigma has density f0f_{0} for all σ\sigma, we find

ℐ︀⁢(σ)=1σ2⁢𝔼⁢((1+Z⁢ℓ0′⁢(Z))2)=cσ2,where ⁢c=∫(1+z⁢ℓ0′⁢(z))2⁢f0⁢(z)⁢dz\mathcal{I}(\sigma)=\frac{1}{\sigma^{2}}\,\mathbb{E}\Bigl{(}\bigl{(}1+Z\ell_{0% }^{\prime}(Z)\bigr{)}^{2}\Bigr{)}=\frac{c}{\sigma^{2}},\qquad\text{where }c=% \int\bigl{(}1+z\ell_{0}^{\prime}(z)\bigr{)}^{2}f_{0}(z)\,\mathrm{d}z

does not depend on σ\sigma.

Finally, the N⁢(0,σ2)N(0,\sigma^{2}) model is the scale model with base density the standard normal density and with ℓ0⁢(z)=−z2/2−log⁡2⁢π\ell_{0}(z)=-z^{2}/2-\log\sqrt{2\pi} and thus ℓ0′⁢(z)=−z\ell_{0}^{\prime}(z)=-z. In this case we have 1+z⁢ℓ0′⁢(z)=1−z21+z\ell_{0}^{\prime}(z)=1-z^{2} and using the substitutions 𝔼⁢(Z2)=1\mathbb{E}(Z^{2})=1 and 𝔼⁢(Z4)=3\mathbb{E}(Z^{4})=3 for the standard normal density ZZ we find

c=𝔼⁢((1−Z2)2)=1−2⁢𝔼⁢(Z2)+𝔼⁢(Z4)=2.c=\mathbb{E}\bigl{(}(1-Z^{2})^{2}\bigr{)}=1-2\mathbb{E}(Z^{2})+\mathbb{E}(Z^{4% })=2.

Thus ℐ︀⁢(σ)=2/σ2\mathcal{I}(\sigma)=2/\sigma^{2}, as claimed. ∎

C.19 Credible Intervals and Bayesian Testing

Solution to exercise 28.1.
  1. 1.
    ​

    Since the posterior density is monotone on each half-line and integrates to one, it converges to zero as θ→±∞\theta\to\pm\infty. The condition ℙ⁢(θ∈Hc|x)=1−α∈(0,1)\mathbb{P}(\theta\in H_{c}\mskip 1.0mu|\mskip 1.0mux)=1-\alpha\in(0,1) excludes c≥π⁢(μ|x)c\geq\pi(\mu\mskip 1.0mu|\mskip 1.0mux), where HcH_{c} contains at most the point μ\mu and has probability zero, and thus 0<c<π⁢(μ|x)0<c<\pi(\mu\mskip 1.0mu|\mskip 1.0mux). On (−∞,μ](-\infty,\mu] the density is continuous and strictly increasing from 0 to π⁢(μ|x)\pi(\mu\mskip 1.0mu|\mskip 1.0mux) and thus there is exactly one l<μl<\mu with π⁢(l|x)=c\pi(l\mskip 1.0mu|\mskip 1.0mux)=c and π⁢(θ|x)≥c\pi(\theta\mskip 1.0mu|\mskip 1.0mux)\geq c for θ≤μ\theta\leq\mu if and only if θ≥l\theta\geq l. Similarly, there is exactly one u>μu>\mu with π⁢(u|x)=c\pi(u\mskip 1.0mu|\mskip 1.0mux)=c and the density is at least cc on [μ,∞)[\mu,\infty) exactly on [μ,u][\mu,u]. Thus we have Hc=[l,u]H_{c}=[l,u] with π⁢(l|x)=c=π⁢(u|x)\pi(l\mskip 1.0mu|\mskip 1.0mux)=c=\pi(u\mskip 1.0mu|\mskip 1.0mux).

  2. 2.
    ​

    Let l=μ−sl=\mu-s and u=μ+tu=\mu+t with s,t>0s,t>0. By symmetry we have π⁢(μ+s|x)=π⁢(μ−s|x)=c=π⁢(μ+t|x)\pi(\mu+s\mskip 1.0mu|\mskip 1.0mux)=\pi(\mu-s\mskip 1.0mu|\mskip 1.0mux)=c=% \pi(\mu+t\mskip 1.0mu|\mskip 1.0mux) and since the density is strictly decreasing on [μ,∞)[\mu,\infty) we find s=ts=t. The two tails (−∞,μ−t)(-\infty,\mu-t) and (μ+t,∞)(\mu+t,\infty) have equal posterior probability by symmetry and together they have probability α\alpha, i.e. each tail has probability α/2\alpha/2. Thus we have l=qα/2l=q_{\alpha/2} and u=q1−α/2u=q_{1-\alpha/2}, i.e. we have found the equal-tailed interval.

  3. 3.
    ​

    The set {θ∣π⁢(θ|x)≥c}\{\theta\mid\pi(\theta\mskip 1.0mu|\mskip 1.0mux)\geq c\} is for 0<c<π⁢(0|x)0<c<\pi(0\mskip 1.0mu|\mskip 1.0mux) the interval [0,u][0,u] with π⁢(u|x)=c\pi(u\mskip 1.0mu|\mskip 1.0mux)=c, obtained by using the same argument as in the first part, but only for the decreasing branch of the density, and using the condition ℙ⁢(θ≤u|x)=1−α\mathbb{P}(\theta\leq u\mskip 1.0mu|\mskip 1.0mux)=1-\alpha to get u=q1−αu=q_{1-\alpha}. The equal-tailed interval [qα/2,q1−α/2][q_{\alpha/2},q_{1-\alpha/2}] is then obtained from [0,q1−α][0,q_{1-\alpha}] by removing the piece [0,qα/2)[0,q_{\alpha/2}) and adding the piece (q1−α,q1−α/2](q_{1-\alpha},q_{1-\alpha/2}], both of which have posterior probability α/2\alpha/2. The density is at least d1=π⁢(qα/2|x)d_{1}=\pi(q_{\alpha/2}\mskip 1.0mu|\mskip 1.0mux) on the first piece and at most d2=π⁢(q1−α|x)d_{2}=\pi(q_{1-\alpha}\mskip 1.0mu|\mskip 1.0mux) on the second piece, where d1>d2d_{1}>d_{2} since the density is strictly decreasing. Writing λ1\lambda_{1} and λ2\lambda_{2} for the lengths of the two pieces we find α/2≥d1⁢λ1\alpha/2\geq d_{1}\lambda_{1} and α/2≤d2⁢λ2\alpha/2\leq d_{2}\lambda_{2} and thus λ1≤α/(2⁢d1)<α/(2⁢d2)≤λ2\lambda_{1}\leq\alpha/(2d_{1})<\alpha/(2d_{2})\leq\lambda_{2}: the added piece is longer than the removed piece and the equal-tailed interval is longer than the HPD interval.

∎

Solution to exercise 28.2.
  1. 1.
    ​

    The equal-tailed 95%95\% credible interval is [0.386,0.861][0.386,0.861], with posterior probability 2.5%2.5\% cut off in each tail. The midpoint of this interval is (0.386+0.861)/2=0.624(0.386+0.861)/2=0.624, which is slightly below the posterior mean 9/14=0.6439/14=0.643.

  2. 2.
    ​

    The posterior density is proportional to θ8⁢(1−θ)4\theta^{8}(1-\theta)^{4} on (0,1)(0,1). Setting the derivative of the logarithm of the posterior density, i.e. 8⁢log⁡θ+4⁢log⁡(1−θ)8\log\theta+4\log(1-\theta), equal to zero, gives the maximum of the density at 8/θ=4/(1−θ)8/\theta=4/(1-\theta) and thus the mode is at θ=8/12=2/3≈0.667\theta=8/12=2/3\approx 0.667. Since the mode is above the mean 0.6430.643, the posterior distribution is skewed.

  3. 3.
    ​

    By definition 28.2 of the HPD interval, the HPD interval contains the values with the highest posterior density and thus is centred around the mode, so that it lies to the right of the equal-tailed interval. The HPD interval is the shortest interval which has posterior probability 0.950.95, in this case of length 0.4730.473 against 0.4750.475. As nn increases, the posterior distribution approaches a symmetric, normal distribution and the mode and mean move together. In the limit, for a symmetric, unimodal posterior, both intervals coincide; this is a result from exercise 28.1.

∎

Solution to exercise 28.3.
  1. 1.
    ​

    Since the prior is flat, we have Beta⁢(1,1)\text{Beta}(1,1) and using n=3n=3 and T=2T=2 from example 26.5 we find the posterior Beta⁢(3,2)\text{Beta}(3,2). Since B⁢(3,2)=Γ⁢(3)⁢Γ⁢(2)/Γ⁢(5)=2/24=1/12B(3,2)=\Gamma(3)\Gamma(2)/\Gamma(5)=2/24=1/12, the density of this posterior is π⁢(θ|x)=12⁢θ2⁢(1−θ)\pi(\theta\mskip 1.0mu|\mskip 1.0mux)=12\theta^{2}(1-\theta) for 0<θ<10<\theta<1.

  2. 2.
    ​

    We can find the posterior probability of H0H_{0} by integrating the density:

    ℙ⁢(θ≤1/2|x)=12⁢∫01/2(θ2−θ3)⁢dθ=12⁢(124−164)=516=0.3125.\mathbb{P}(\theta\leq 1/2\mskip 1.0mu|\mskip 1.0mux)=12\int_{0}^{1/2}(\theta^{% 2}-\theta^{3})\,\mathrm{d}\theta=12\Bigl{(}\frac{1}{24}-\frac{1}{64}\Bigr{)}=% \frac{5}{16}=0.3125.

    Thus we have ℙ⁢(H1|x)=11/16\mathbb{P}(H_{1}\mskip 1.0mu|\mskip 1.0mux)=11/16 and the posterior odds for H0H_{0} against H1H_{1} are 5/11≈0.455/11\approx 0.45.

  3. 3.
    ​

    The data have multiplied the odds by 5/115/11. Using Bayes’ theorem for the events H0H_{0} and H1H_{1}, we find that the posterior probability of HjH_{j} is ℙ⁢(Hj)\mathbb{P}(H_{j}) times the average of the likelihoods over HjH_{j}, with respect to the prior conditioned on HjH_{j}, divided by the marginal density of the data. The factor by which the odds change is the ratio of these two averaged likelihoods. In this case, the likelihood is θ2⁢(1−θ)\theta^{2}(1-\theta) and the conditional priors are uniform on (0,1/2)(0,1/2) and on (1/2,1)(1/2,1), and thus the factor is

    2⁢∫01/2θ2⁢(1−θ)⁢dθ2⁢∫1/21θ2⁢(1−θ)⁢dθ=5/1921/12−5/192=511.\frac{2\int_{0}^{1/2}\theta^{2}(1-\theta)\,\mathrm{d}\theta}{2\int_{1/2}^{1}% \theta^{2}(1-\theta)\,\mathrm{d}\theta}=\frac{5/192}{1/12-5/192}=\frac{5}{11}.

    This agrees with the result from the previous part. Note that this is not the ratio of likelihoods at two values, since each hypothesis includes a range of parameter values, and these need to be weighted by the prior before the two hypotheses can be compared. Only for two simple hypotheses do the averages coincide with single likelihood values, as in case (28.3).

  4. 4.
    ​

    The bounds ll and uu satisfy 12⁢∫0l(θ2−θ3)⁢dθ=0.02512\int_{0}^{l}(\theta^{2}-\theta^{3})\,\mathrm{d}\theta=0.025 and 12⁢∫0u(θ2−θ3)⁢dθ=0.97512\int_{0}^{u}(\theta^{2}-\theta^{3})\,\mathrm{d}\theta=0.975, i.e. we have

    4⁢l3−3⁢l4=0.025and4⁢u3−3⁢u4=0.975.4l^{3}-3l^{4}=0.025\qquad\text{and}\qquad 4u^{3}-3u^{4}=0.975.

    These are polynomial equations of degree four and while they can be solved explicitly, the solutions are cumbersome and not very instructive; for a general Beta posterior, the bounds will not be algebraic. The quantiles are best found numerically, for example using the R function qbeta to get l=0.194l=0.194 and u=0.932u=0.932 and thus the 95%95\% equal-tailed credible interval is [0.194,0.932][0.194,0.932].

∎

Solution to exercise 28.4.
  1. 1.
    ​

    Using the formula from example 28.7, with n=25n=25, σ2=4\sigma^{2}=4, θ0=10\theta_{0}=10, θ1=11\theta_{1}=11 and x¯=10.9\bar{x}=10.9, we find that the Bayes factor in favour of H0H_{0} is

    f⁢(x|θ0)f⁢(x|θ1)=exp⁡(25⋅1⋅(21−21.8)8)=e−2.5=0.082.\frac{f(x\mskip 1.0mu|\mskip 1.0mu\theta_{0})}{f(x\mskip 1.0mu|\mskip 1.0mu% \theta_{1})}=\exp\Bigl{(}\frac{25\cdot 1\cdot(21-21.8)}{8}\Bigr{)}=e^{-2.5}=0.% 082.

    The data give evidence in favour of H1H_{1} over H0H_{0} by a factor of e2.5≈12.2e^{2.5}\approx 12.2.

  2. 2.
    ​

    The prior odds are 33 and thus the posterior odds are 3⋅0.082=0.2463\cdot 0.082=0.246 and ℙ⁢(H0|x)=0.246/1.246=0.198\mathbb{P}(H_{0}\mskip 1.0mu|\mskip 1.0mux)=0.246/1.246=0.198. The data have changed the prior opinion, which was three to one in favour of H0H_{0}, to a posterior opinion of approximately four to one against it.

  3. 3.
    ​

    The pp-value 0.01220.0122 is the probability, under θ=10\theta=10, of obtaining a sample mean of at least 10.910.9. This value refers to the data, includes outcomes more extreme than the one observed and makes no reference to any alternative. In contrast, the posterior probability 0.1980.198 is a probability for the hypothesis; it uses only the observed data, compares the likelihoods at exactly 1010 and 1111 and takes into account the prior odds, which were chosen in favour of H0H_{0}. Even if the prior odds had been 11, the posterior probability would have been 0.082/1.082=0.0760.082/1.082=0.076, or six times the pp-value. Both numbers show that the data are evidence against H0H_{0}, but in different ways. Neither number is a rewording of the other.

  4. 4.
    ​

    The posterior odds are the prior odds times e−2.5e^{-2.5}, and the posterior odds are greater than 11 if and only if the prior odds are greater than e2.5≈12.2e^{2.5}\approx 12.2. Only a prior which assumed H0H_{0} to be more likely than H1H_{1}, by a factor of more than twelve, would result in H0H_{0} being the more likely hypothesis after these data are observed.

∎

Solution to exercise 28.5.
  1. 1.
    ​

    By (28.1) we have 1−w=(n/σ2)/(1/τ2+n/σ2)=(n/σ2)⁢τn21-w=(n/\sigma^{2})/(1/\tau^{2}+n/\sigma^{2})=(n/\sigma^{2})\,\tau_{n}^{2}, and thus τn2=(1−w)⁢σ2/n\tau_{n}^{2}=(1-w)\sigma^{2}/n. The posterior mean is mn=w⁢m+(1−w)⁢X¯m_{n}=wm+(1-w)\bar{X}, and thus mn−θ=w⁢(m−θ)+(1−w)⁢(X¯−θ)m_{n}-\theta=w(m-\theta)+(1-w)(\bar{X}-\theta). Under ℙθ\mathbb{P}_{\theta} we have X¯−θ∼N⁢(0,σ2/n)\bar{X}-\theta\sim N(0,\sigma^{2}/n) by proposition A.3, and an affine function of a normal variable is normal, so that mn−θ∼N⁢(w⁢(m−θ),(1−w)2⁢σ2/n)m_{n}-\theta\sim N\bigl{(}w(m-\theta),(1-w)^{2}\sigma^{2}/n\bigr{)}.

  2. 2.
    ​

    Write z=z1−α/2z=z_{1-\alpha/2} and δ=n⁢(m−θ)/σ\delta=\sqrt{n}\,(m-\theta)/\sigma for the standardised distance between the prior mean and the true value. The interval covers θ\theta if and only if |mn−θ|≤z⁢τn=z⁢1−w⁢σ/n|m_{n}-\theta|\leq z\tau_{n}=z\sqrt{1-w}\,\sigma/\sqrt{n}. Subtracting the mean w⁢(m−θ)w(m-\theta) and dividing by the standard deviation (1−w)⁢σ/n(1-w)\sigma/\sqrt{n} turns mn−θm_{n}-\theta into a standard normal variable, and the bounds into (±z⁢1−w−w⁢δ)/(1−w)(\pm z\sqrt{1-w}-w\delta)/(1-w). Thus the coverage probability is

    ℙθ⁢(θ∈interval)=Φ⁢(z⁢1−w−w⁢δ1−w)−Φ⁢(−z⁢1−w−w⁢δ1−w).\mathbb{P}_{\theta}\bigl{(}\theta\in\text{interval}\bigr{)}=\Phi\Bigl{(}\frac{% z\sqrt{1-w}-w\delta}{1-w}\Bigr{)}-\Phi\Bigl{(}\frac{-z\sqrt{1-w}-w\delta}{1-w}% \Bigr{)}.

    It depends on θ\theta only through δ\delta, and it is symmetric in δ\delta.

  3. 3.
    ​

    Here 1/τ2=11/\tau^{2}=1 and n/σ2=4n/\sigma^{2}=4, and thus w=0.2w=0.2, 1−w=0.81-w=0.8, 1−w=0.894\sqrt{1-w}=0.894 and z⁢1−w=1.753z\sqrt{1-w}=1.753. At θ=m=12\theta=m=12 we have δ=0\delta=0 and the coverage probability is 2⁢Φ⁢(1.753/0.8)−1=2⁢Φ⁢(2.19)−1=0.9712\Phi(1.753/0.8)-1=2\Phi(2.19)-1=0.971, above the nominal 0.950.95. At θ=10\theta=10 we have δ=4⋅2/2=4\delta=4\cdot 2/2=4 and w⁢δ=0.8w\delta=0.8, and the coverage probability is Φ⁢(0.953/0.8)−Φ⁢(−2.553/0.8)=Φ⁢(1.19)−Φ⁢(−3.19)=0.883−0.001=0.882\Phi(0.953/0.8)-\Phi(-2.553/0.8)=\Phi(1.19)-\Phi(-3.19)=0.883-0.001=0.882, below the nominal value. As |θ−m|→∞|\theta-m|\to\infty we have |δ|→∞|\delta|\to\infty, both arguments of Φ\Phi tend to the same infinity and the coverage probability tends to zero: the interval is pulled towards the prior mean by a fixed fraction ww of the distance, and once that pull exceeds the half-width the interval almost never contains the truth.

  4. 4.
    ​

    It is not a confidence interval with level 0.950.95, because definition 25.1 requires coverage probability at least 0.950.95 for every θ\theta, and the previous part exhibits values of θ\theta where it is smaller. The credible interval trades the uniform guarantee for better behaviour near the prior mean, which is where a well-chosen prior expects the truth to be. As n→∞n\to\infty with θ\theta fixed we have w≈σ2/(n⁢τ2)→0w\approx\sigma^{2}/(n\tau^{2})\to 0, and thus w⁢δ≈σ⁢(m−θ)/(τ2⁢n)→0w\delta\approx\sigma(m-\theta)/(\tau^{2}\sqrt{n})\to 0 and 1−w→1\sqrt{1-w}\to 1, so that the coverage probability tends to 2⁢Φ⁢(z)−1=1−α2\Phi(z)-1=1-\alpha for every θ\theta. This is the large-sample agreement of section 28.3 seen from the frequentist side.

∎

C.20 Standard Tests Revisited

Solution to exercise 29.1.
  1. 1.
    ​

    We have the sum of observations 154154 and thus x¯=154/8=19.25\bar{x}=154/8=19.25. The sum of squared deviations from x¯\bar{x} equals 3.303.30 and thus we have s2=3.30/7=0.471s^{2}=3.30/7=0.471 and s=0.687s=0.687. The observed value of the test statistic is

    t=8⁢(19.25−20)0.687=2.828⋅(−0.75)0.687=−3.09.t=\frac{\sqrt{8}\,(19.25-20)}{0.687}=\frac{2.828\cdot(-0.75)}{0.687}=-3.09.
  2. 2.
    ​

    From proposition 29.1 we know that the test against μ<20\mu<20 rejects if t<−t7,1−αt<-t_{7,1-\alpha}. For level 0.050.05 the critical value is −t7,0.95=−1.895-t_{7,0.95}=-1.895 and for level 0.010.01 the critical value is −t7,0.99=−2.998-t_{7,0.99}=-2.998; since −3.09-3.09 lies below both of these critical values, the hypothesis H0H_{0} is rejected at both levels. The pp-value is given by ℙ⁢(T≤−3.09)=ℙ⁢(T≥3.09)\mathbb{P}(T\leq-3.09)=\mathbb{P}(T\geq 3.09) for T∼t7T\sim t_{7}, by symmetry, and since 3.093.09 is between t7,0.99=2.998t_{7,0.99}=2.998 and t7,0.995=3.499t_{7,0.995}=3.499, the value is between 0.0050.005 and 0.010.01.

  3. 3.
    ​

    The two-sided test rejects if |t|>t7,1−α/2|t|>t_{7,1-\alpha/2}. For level 0.050.05 we have |t|=3.09>t7,0.975=2.365|t|=3.09>t_{7,0.975}=2.365 and thus H0H_{0} is rejected. For level 0.010.01 we have 3.09<t7,0.995=3.4993.09<t_{7,0.995}=3.499 and thus H0H_{0} is not rejected. The two-sided pp-value is 2⁢ℙ⁢(T≥3.09)2\,\mathbb{P}(T\geq 3.09), which is twice the value of the one-sided test, so it is between 0.010.01 and 0.020.02.

  4. 4.
    ​

    The 99%99\% interval is x¯±t7,0.995⁢s/8=19.25±3.499⋅0.687/2.828=19.25±0.849\bar{x}\pm t_{7,0.995}\,s/\sqrt{8}=19.25\pm 3.499\cdot 0.687/2.828=19.25\pm 0.% 849. This interval is [18.40,20.10][18.40,20.10]. The interval contains the value 2020 and by theorem 25.6 this is equivalent to the non-rejection of H0:μ=20H_{0}\colon\mu=20 by the two-sided test with level 0.010.01 in the previous part. The 95%95\% interval 19.25±2.365⋅0.243=[18.68,19.82]19.25\pm 2.365\cdot 0.243=[18.68,19.82] does not contain 2020 and thus the rejection at level 0.050.05 is consistent with the results of the other parts.

∎

Solution to exercise 29.2.
  1. 1.
    ​

    The value of the test statistic (29.2) is v=14⋅2.9/1.5=27.07v=14\cdot 2.9/1.5=27.07. We can compare this value to quantiles of the χ142\chi^{2}_{14} distribution: at level 0.050.05 the critical value is χ14,0.952=23.685\chi^{2}_{14,0.95}=23.685 and since 27.07>23.68527.07>23.685 we can reject H0H_{0}, whereas at level 0.010.01 the critical value is χ14,0.992=29.141\chi^{2}_{14,0.99}=29.141 and since 27.07<29.14127.07<29.141 we cannot reject H0H_{0}. The pp-value, ℙ⁢(W≥27.07)\mathbb{P}(W\geq 27.07) for W∼χ142W\sim\chi^{2}_{14}, is between 0.010.01 and 0.0250.025, since vv is between χ14,0.9752=26.119\chi^{2}_{14,0.975}=26.119 and χ14,0.992=29.141\chi^{2}_{14,0.99}=29.141.

  2. 2.
    ​

    The two-sided test with level 0.050.05 does not reject, if

    v∈[χ14,0.0252,χ14,0.9752]=[5.629,26.119].v\in[\chi^{2}_{14,0.025},\chi^{2}_{14,0.975}]=[5.629,26.119].

    Since 27.07>26.11927.07>26.119, we reject H0H_{0}.

  3. 3.
    ​

    Let c=χn−1,1−α2c=\chi^{2}_{n-1,1-\alpha}. Assuming the true variance is σ2\sigma^{2}, the value W=(n−1)⁢S2/σ2W=(n-1)S^{2}/\sigma^{2} is χn−12\chi^{2}_{n-1} distributed by proposition A.3 and thus V=(n−1)⁢S2/σ02=W⁢σ2/σ02V=(n-1)S^{2}/\sigma_{0}^{2}=W\sigma^{2}/\sigma_{0}^{2}. We have

    β⁢(σ2)=ℙσ2⁢(V>c)=ℙ⁢(W>c⁢σ02σ2).\beta(\sigma^{2})=\mathbb{P}_{\sigma^{2}}(V>c)=\mathbb{P}\Bigl{(}W>\frac{c\,% \sigma_{0}^{2}}{\sigma^{2}}\Bigr{)}.

    As σ2\sigma^{2} increases, the threshold c⁢σ02/σ2c\sigma_{0}^{2}/\sigma^{2} decreases and the probability of WW exceeding a smaller number increases, i.e. β\beta increases as a function of σ2\sigma^{2}. By definition 20.5, the size of the test for the composite null hypothesis σ2≤σ02\sigma^{2}\leq\sigma_{0}^{2} is supσ2≤σ02β⁢(σ2)=β⁢(σ02)=ℙ⁢(W>c)=α\sup_{\sigma^{2}\leq\sigma_{0}^{2}}\beta(\sigma^{2})=\beta(\sigma_{0}^{2})=% \mathbb{P}(W>c)=\alpha, where the bound is achieved at the boundary.

  4. 4.
    ​

    The two-sided test with level α\alpha rejects large values of VV only if V>χn−1,1−α/22V>\chi^{2}_{n-1,1-\alpha/2}. Since this is a larger threshold than the χn−1,1−α2\chi^{2}_{n-1,1-\alpha} of the one-sided test, the power of the two-sided test against the alternative σ2>σ02\sigma^{2}>\sigma_{0}^{2} is smaller than the power of the one-sided test. An observed value between the two thresholds is rejected by the one-sided test but not by the two-sided test. The two-sided test has power above α\alpha against the alternative σ2<σ02\sigma^{2}<\sigma_{0}^{2}, where the one-sided test has probability of rejection below α\alpha. For the given data, both tests reject the hypothesis at level 0.050.05, but at level 0.0250.025 the one-sided test, with threshold χ14,0.9752=26.119\chi^{2}_{14,0.975}=26.119, still rejects, whereas the two-sided test compares 27.0727.07 to χ14,0.98752=28.42\chi^{2}_{14,0.9875}=28.42 and does not reject.

∎

Solution to exercise 29.3.
  1. 1.
    ​

    We can describe the differences Di=Xi−YiD_{i}=X_{i}-Y_{i} using i.i.d. N⁢(μD,σD2)N(\mu_{D},\sigma_{D}^{2}), where both parameters are unknown. We test H0:μD=0H_{0}\colon\mu_{D}=0 against H1:μD>0H_{1}\colon\mu_{D}>0, since we expect the reaction time to be smaller after training. The differences are 4,−1,6,3,6,54,-1,6,3,6,5 and the sum of these differences is 2323, i.e. we have d¯=3.833\bar{d}=3.833. The sum of squared deviations from d¯\bar{d} is 34.8334.83, and thus we have sD2=34.83/5=6.967s_{D}^{2}=34.83/5=6.967 and sD=2.639s_{D}=2.639. The test statistic from proposition 29.1, for μ0=0\mu_{0}=0, is

    t=6⋅3.8332.639=3.56,t=\frac{\sqrt{6}\cdot 3.833}{2.639}=3.56,

    where 55 is the number of degrees of freedom. Since 3.56>t5,0.95=2.0153.56>t_{5,0.95}=2.015 we can reject the hypothesis at level 0.050.05 and since 3.56>t5,0.99=3.3653.56>t_{5,0.99}=3.365 we can also reject the hypothesis at level 0.010.01: the data show a reduction in reaction time.

  2. 2.
    ​

    The 95%95\% tt-interval from example 25.4 for μD\mu_{D} is d¯±t5,0.975⁢sD/6=3.833±2.571⋅2.639/2.449=3.833±2.770\bar{d}\pm t_{5,0.975}\,s_{D}/\sqrt{6}=3.833\pm 2.571\cdot 2.639/2.449=3.833% \pm 2.770, which is [1.06,6.60][1.06,6.60].

  3. 3.
    ​

    The row means are x¯=419/6=69.83\bar{x}=419/6=69.83 and y¯=396/6=66.00\bar{y}=396/6=66.00, the pooled variance is sp2=(5⋅59.77+5⋅42.80)/10=51.28s_{p}^{2}=(5\cdot 59.77+5\cdot 42.80)/10=51.28 and thus we have sp=7.161s_{p}=7.161. The test statistic (29.3) is

    t=69.83−66.007.161⁢1/6+1/6=3.8334.135=0.93,t=\frac{69.83-66.00}{7.161\sqrt{1/6+1/6}}=\frac{3.833}{4.135}=0.93,

    where 1010 is the number of degrees of freedom. Since 0.93<t10,0.95=1.8120.93<t_{10,0.95}=1.812 the colleague does not reject the hypothesis at level 0.050.05. The two-sample test is wrong, because the two measurements for one subject are not independent (a slow person is likely to be slow after training, too), whereas the two-sample test assumes that the twelve observations are independent. The difference in opinion is caused by the variability of the data. Between subjects the standard deviation of the reaction times is about 77. The two-sample test compares the mean difference 3.833.83 to this spread. Inside subjects the differences have standard deviation only 2.642.64. The paired test compares 3.833.83 to this smaller spread. By taking differences we remove the effect of the subjects from the comparison, and this is the idea of the paired test.

∎

Solution to exercise 29.4.
  1. 1.
    ​

    For μ0=9\mu_{0}=9 the value of the test statistic (29.1) is t=4⁢(10.3−9)/2.1=2.48t=4\,(10.3-9)/2.1=2.48 and, since |t|>t15,0.975=2.131|t|>t_{15,0.975}=2.131, the two-sided test rejects at level 0.050.05, as found using duality in exercise 25.4. For μ0=11.5\mu_{0}=11.5 we have t=4⁢(10.3−11.5)/2.1=−2.29t=4\,(10.3-11.5)/2.1=-2.29 and |t|=2.29>2.131|t|=2.29>2.131, so H0H_{0} is rejected again. This agrees with the interval, since 11.511.5 is outside the interval [9.18,11.42][9.18,11.42].

  2. 2.
    ​

    From theorem 25.6 we know that the two-sided test from proposition 29.3 rejects σ02\sigma_{0}^{2} at level 0.050.05 if and only if σ02\sigma_{0}^{2} is outside the 95%95\% confidence interval [2.41,10.56][2.41,10.56] from example 25.5. This is the inverse of the test, so σ02=2\sigma_{0}^{2}=2 is rejected and σ02=9\sigma_{0}^{2}=9 is not. Looking at the values directly, for (n−1)⁢s2=15⋅4.41=66.15(n-1)s^{2}=15\cdot 4.41=66.15 the test statistic corresponding to σ02=2\sigma_{0}^{2}=2 is v=66.15/2=33.08v=66.15/2=33.08, which is greater than χ15,0.9752=27.488\chi^{2}_{15,0.975}=27.488, so H0H_{0} is rejected. For σ02=9\sigma_{0}^{2}=9 the test statistic v=66.15/9=7.35v=66.15/9=7.35 falls into the interval [6.262,27.488][6.262,27.488] and thus H0H_{0} is not rejected.

  3. 3.
    ​

    For μ0=9.5\mu_{0}=9.5 the test statistic is t=4⁢(10.3−9.5)/2.1=1.52t=4\,(10.3-9.5)/2.1=1.52 and, since 1.52≤t15,0.95=1.7531.52\leq t_{15,0.95}=1.753, the one-sided test does not reject at level 0.050.05. The two-sided interval is the inverse of the two-sided test with level α\alpha, i.e. the upper tail of the two-sided test has probability α/2\alpha/2. If we decided a one-sided test by checking whether μ0\mu_{0} is less than the lower bound of the interval, we would obtain the one-sided test at level α/2=0.025\alpha/2=0.025, not at level 0.050.05. Thus, the one-sided test with 0.050.05 is not the same as the two-sided test. The confidence set corresponding to the one-sided test with 0.050.05 is, as in exercise 25.4, the one-sided interval [x¯−t15,0.95⁢s/n,∞)=[10.3−1.753⋅0.525,∞)=[9.38,∞)\bigl{[}\bar{x}-t_{15,0.95}\,s/\sqrt{n},\ \infty\bigr{)}=[10.3-1.753\cdot 0.52% 5,\ \infty)=[9.38,\infty), and 9.59.5 is contained in this interval, as expected.

∎

C.21 The Exponential Model: From Start to Finish

Here we solve the exercises at the end of lecture 31.

Solution to exercise 31.1.
  1. 1.
    ​

    The log-likelihood function is ℓ⁢(θ)=n⁢log⁡θ−θ⁢T\ell(\theta)=n\log\theta-\theta T, with T=∑ixi=n⁢x¯T=\sum_{i}x_{i}=n\bar{x}, and the MLE is θ^=n/T\hat{\theta}=n/T. Using θ^⁢T=n\hat{\theta}T=n and θ0⁢T=n⁢θ0/θ^\theta_{0}T=n\theta_{0}/\hat{\theta} we find

    W=2⁢(ℓ⁢(θ^)−ℓ⁢(θ0))=2⁢(n⁢log⁡θ^−n−n⁢log⁡θ0+n⁢θ0/θ^)=2⁢n⁢(log⁡θ^θ0−1+θ0θ^).W=2\bigl{(}\ell(\hat{\theta})-\ell(\theta_{0})\bigr{)}=2\bigl{(}n\log\hat{% \theta}-n-n\log\theta_{0}+n\theta_{0}/\hat{\theta}\bigr{)}=2n\Bigl{(}\log\frac% {\hat{\theta}}{\theta_{0}}-1+\frac{\theta_{0}}{\hat{\theta}}\Bigr{)}.

    Using u=θ0⁢x¯=θ0/θ^u=\theta_{0}\bar{x}=\theta_{0}/\hat{\theta} we have log⁡(θ^/θ0)=−log⁡u\log(\hat{\theta}/\theta_{0})=-\log u and θ0/θ^=u\theta_{0}/\hat{\theta}=u, and thus W=2⁢n⁢(u−1−log⁡u)W=2n(u-1-\log u).

  2. 2.
    ​

    Let g⁢(u)=u−1−log⁡ug(u)=u-1-\log u for u>0u>0. Then g⁢(1)=0g(1)=0 and g′⁢(u)=1−1/ug^{\prime}(u)=1-1/u. Since this value is negative for u<1u<1 and positive for u>1u>1, u=1u=1 must be the unique minimum of gg and g⁢(u)>0g(u)>0 for u≠1u\neq 1. Since W=2⁢n⁢g⁢(u)W=2ng(u), we have W≥0W\geq 0, with equality if and only if u=1u=1, i.e. if and only if θ^=θ0\hat{\theta}=\theta_{0}.

  3. 3.
    ​

    For θ0=0.5\theta_{0}=0.5 we have u=1.25u=1.25 and W=20⁢(1.25−1−log⁡1.25)=20⋅0.0269=0.54W=20\,(1.25-1-\log 1.25)=20\cdot 0.0269=0.54. This value is less than χ1,0.952=3.84\chi^{2}_{1,0.95}=3.84, so H0:θ=0.5H_{0}\colon\theta=0.5 is not rejected at level 0.050.05. For θ0=1\theta_{0}=1 we have u=2.5u=2.5 and W=20⁢(2.5−1−log⁡2.5)=20⋅0.584=11.67W=20\,(2.5-1-\log 2.5)=20\cdot 0.584=11.67. This value is greater than 3.843.84, so H0:θ=1H_{0}\colon\theta=1 is rejected: the data are not compatible with a mean lifetime of one year.

This completes the solution. ∎

Solution to exercise 31.2.
  1. 1.
    ​

    From example 11.8 we know that ℐ︀⁢(θ)=1/θ2\mathcal{I}(\theta)=1/\theta^{2}. Thus, the Jeffreys prior is πJ⁢(θ)∝ℐ︀⁢(θ)=1/θ\pi_{J}(\theta)\propto\sqrt{\mathcal{I}(\theta)}=1/\theta. Since ∫01θ−1⁢dθ=∞\int_{0}^{1}\theta^{-1}\,\mathrm{d}\theta=\infty, the prior has infinite integral and thus is improper.

  2. 2.
    ​

    The posterior density is proportional to θ−1⋅θn⁢e−θ⁢T=θn−1⁢e−θ⁢T\theta^{-1}\cdot\theta^{n}e^{-\theta T}=\theta^{n-1}e^{-\theta T}. This is the Gamma(n,T)(n,T) density up to the constant Tn/Γ⁢(n)T^{n}/\Gamma(n). Since T>0T>0, this constant is finite, so the posterior is proper and corresponds to the Gamma(n,T)(n,T) distribution.

  3. 3.
    ​

    From exercise 11.5 we know that the information about the mean is ℐ︀μ⁢(μ)=1/μ2\mathcal{I}_{\mu}(\mu)=1/\mu^{2}. Thus, the Jeffreys prior for μ\mu is πJ⁢(μ)∝1/μ\pi_{J}(\mu)\propto 1/\mu. Using the transformation πJ⁢(θ)\pi_{J}(\theta) with θ=1/μ\theta=1/\mu and |d⁢θ/d⁢μ|=1/μ2|\mathrm{d}\theta/\mathrm{d}\mu|=1/\mu^{2} we find the transformed density πJ⁢(1/μ)/μ2∝μ/μ2=1/μ\pi_{J}(1/\mu)/\mu^{2}\propto\mu/\mu^{2}=1/\mu. This is the same prior as πJ⁢(μ)\pi_{J}(\mu) above, as expected from proposition 26.9.

  4. 4.
    ​

    For θ∼Gamma⁢(n,T)\theta\sim\text{Gamma}(n,T) we have 2⁢T⁢θ∼Gamma⁢(n,1/2)=χ2⁢n22T\theta\sim\text{Gamma}(n,1/2)=\chi^{2}_{2n} and thus the equal-tailed 95%95\% credible interval is [χ2⁢n,0.0252/(2⁢T),χ2⁢n,0.9752/(2⁢T)][\chi^{2}_{2n,0.025}/(2T),\ \chi^{2}_{2n,0.975}/(2T)]. These are the same numbers as for the exact confidence interval. The confidence interval covers the fixed value θ\theta in 95%95\% of repeated samples. The credible interval, on the other hand, covers the random variable θ\theta with posterior probability 0.950.95 given the observed data.

This completes the solution. ∎

Solution to exercise 31.3.
  1. (a)
    ​

    The distribution function of one observation is F⁢(y)=y/θF(y)=y/\theta for 0≤y≤θ0\leq y\leq\theta, with density f⁢(y)=1/θf(y)=1/\theta there. By proposition 8.9 the minimum has density

    f(1)⁢(y)=n⁢(1−F⁢(y))n−1⁢f⁢(y)=nθ⁢(1−yθ)n−1,0≤y≤θ,f_{(1)}(y)=n\bigl{(}1-F(y)\bigr{)}^{n-1}f(y)=\frac{n}{\theta}\Bigl{(}1-\frac{y% }{\theta}\Bigr{)}^{n-1},\qquad 0\leq y\leq\theta,

    and zero otherwise. Substituting u=y/θu=y/\theta we find

    𝔼θ⁢(X(1))=θ⁢∫01n⁢u⁢(1−u)n−1⁢du=θ⁢n⁢B⁢(2,n)=θ⁢n⁢1!⁢(n−1)!(n+1)!=θn+1,\mathbb{E}_{\theta}(X_{(1)})=\theta\int_{0}^{1}n\,u(1-u)^{n-1}\,\mathrm{d}u=% \theta\,n\,B(2,n)=\theta\,n\,\frac{1!\,(n-1)!}{(n+1)!}=\frac{\theta}{n+1},

    using the beta function of table A.2, or equivalently the mean 1/(n+1)1/(n+1) of the Beta⁢(1,n)\text{Beta}(1,n) distribution of X(1)/θX_{(1)}/\theta.

  2. (b)
    ​

    By part (a) and the linearity of expectation, 𝔼θ⁢(U)=(n+1)⁢θ/(n+1)=θ\mathbb{E}_{\theta}(U)=(n+1)\,\theta/(n+1)=\theta for all θ\theta, and thus UU is unbiased. The same substitution gives

    𝔼θ⁢(X(1)2)\displaystyle\mathbb{E}_{\theta}(X_{(1)}^{2})
    =θ2⁢∫01n⁢u2⁢(1−u)n−1⁢du=θ2⁢n⁢B⁢(3,n)\displaystyle=\theta^{2}\int_{0}^{1}n\,u^{2}(1-u)^{n-1}\,\mathrm{d}u=\theta^{2% }\,n\,B(3,n)
    =θ2⁢2⁢n⁢(n−1)!(n+2)!=2⁢θ2(n+1)⁢(n+2),\displaystyle=\theta^{2}\,\frac{2n\,(n-1)!}{(n+2)!}=\frac{2\theta^{2}}{(n+1)(n% +2)},

    and thus

    Varθ(X(1))\displaystyle\mathop{\mathrm{Var}}\nolimits_{\theta}(X_{(1)})
    =2⁢θ2(n+1)⁢(n+2)−θ2(n+1)2\displaystyle=\frac{2\theta^{2}}{(n+1)(n+2)}-\frac{\theta^{2}}{(n+1)^{2}}
    =θ2⁢(2⁢(n+1)−(n+2))(n+1)2⁢(n+2)=n⁢θ2(n+1)2⁢(n+2)\displaystyle=\frac{\theta^{2}\bigl{(}2(n+1)-(n+2)\bigr{)}}{(n+1)^{2}(n+2)}=% \frac{n\,\theta^{2}}{(n+1)^{2}(n+2)}

    and Varθ(U)=(n+1)2⁢Varθ(X(1))=n⁢θ2/(n+2)\mathop{\mathrm{Var}}\nolimits_{\theta}(U)=(n+1)^{2}\mathop{\mathrm{Var}}% \nolimits_{\theta}(X_{(1)})=n\theta^{2}/(n+2). Since UU is unbiased, MSE(U)=Varθ(U)→θ2\mathop{\mathrm{MSE}}\nolimits(U)=\mathop{\mathrm{Var}}\nolimits_{\theta}(U)% \to\theta^{2} as n→∞n\to\infty, and thus the mean squared error does not tend to zero. This alone does not disprove consistency, since theorem 2.10 is only a sufficient condition; but directly, for 0<ε<θ0<\varepsilon<\theta we have

    ℙθ⁢(|U−θ|>ε)\displaystyle\mathbb{P}_{\theta}\bigl{(}|U-\theta|>\varepsilon\bigr{)}
    ≥ℙθ⁢(U<θ−ε)=ℙθ⁢(X(1)<θ−εn+1)\displaystyle\geq\mathbb{P}_{\theta}(U<\theta-\varepsilon)=\mathbb{P}_{\theta}% \Bigl{(}X_{(1)}<\frac{\theta-\varepsilon}{n+1}\Bigr{)}
    =1−(1−θ−ε(n+1)⁢θ)n→1−e−(θ−ε)/θ>0,\displaystyle=1-\Bigl{(}1-\frac{\theta-\varepsilon}{(n+1)\theta}\Bigr{)}^{n}% \to 1-e^{-(\theta-\varepsilon)/\theta}>0,

    and thus UU does not converge to θ\theta in probability: UU is not consistent in the sense of definition 2.8. The reason is that UU uses only the smallest observation, which carries little information about the upper endpoint θ\theta.

  3. (c)
    ​

    Writing the density of one observation as θ−1⁢𝟏{0≤x≤θ}\theta^{-1}\mathbf{1}_{\{0\leq x\leq\theta\}}, the joint density is

    f⁢(x;θ)=θ−n⁢∏i=1n𝟏{0≤xi≤θ}=θ−n⁢ 1{x(n)≤θ}⋅𝟏{x(1)≥0},f(x;\theta)=\theta^{-n}\prod_{i=1}^{n}\mathbf{1}_{\{0\leq x_{i}\leq\theta\}}=% \theta^{-n}\,\mathbf{1}_{\{x_{(n)}\leq\theta\}}\cdot\mathbf{1}_{\{x_{(1)}\geq 0% \}},

    since all xix_{i} are at most θ\theta if and only if the largest is, and all are non-negative if and only if the smallest is. This is of the form (8.1) with g⁢(x(n);θ)=θ−n⁢𝟏{x(n)≤θ}g(x_{(n)};\theta)=\theta^{-n}\mathbf{1}_{\{x_{(n)}\leq\theta\}} and h⁢(x)=𝟏{x(1)≥0}h(x)=\mathbf{1}_{\{x_{(1)}\geq 0\}}, and thus X(n)X_{(n)} is sufficient for θ\theta by theorem 8.2. The indicator involving θ\theta has to stay in gg.

  4. (d)
    ​

    Conditionally on X(n)=mX_{(n)}=m, the other n−1n-1 observations are i.i.d. uniform on (0,m)(0,m), and X(1)X_{(1)} is their minimum. By part (a) with n−1n-1 in place of nn and mm in place of θ\theta, the minimum of n−1n-1 i.i.d. uniform observations on (0,m)(0,m) has mean m/nm/n, and thus

    𝔼θ⁢(X(1)|X(n)=m)=mn,V=𝔼θ⁢(U|X(n))=(n+1)⁢X(n)n.\mathbb{E}_{\theta}(X_{(1)}\mskip 1.0mu|\mskip 1.0muX_{(n)}=m)=\frac{m}{n},% \qquad V=\mathbb{E}_{\theta}(U\mskip 1.0mu|\mskip 1.0muX_{(n)})=(n+1)\,\frac{X% _{(n)}}{n}.

    The Rao–Blackwell theorem, theorem 10.1, applies because X(n)X_{(n)} is sufficient by part (c) and UU is unbiased with finite variance by part (b). It guarantees that VV does not depend on θ\theta, which the formula confirms, that VV is unbiased for θ\theta, and that Varθ(V)≤Varθ(U)\mathop{\mathrm{Var}}\nolimits_{\theta}(V)\leq\mathop{\mathrm{Var}}\nolimits_{% \theta}(U).

  5. (e)
    ​

    By proposition 10.5 the maximum X(n)X_{(n)} is complete, and it is sufficient by part (c). Since VV is a function of X(n)X_{(n)} and unbiased for θ\theta, the Lehmann–Scheffe theorem, theorem 10.4, shows that VV is the UMVUE for θ\theta. Its variance θ2/(n⁢(n+2))\theta^{2}/\bigl{(}n(n+2)\bigr{)} is smaller than Varθ(U)=n⁢θ2/(n+2)\mathop{\mathrm{Var}}\nolimits_{\theta}(U)=n\theta^{2}/(n+2) by a factor n2n^{2}, and it lies below the Cramer–Rao value θ2/n\theta^{2}/n. There is no contradiction, because theorem 13.2 does not apply to this model: its conditions, section 11.1, require the support of the density not to depend on θ\theta, and here the support is (0,θ)(0,\theta). The identity Covθ(W,ℓ′⁢(θ))=1\mathop{\mathrm{Cov}}\nolimits_{\theta}\bigl{(}W,\ell^{\prime}(\theta)\bigr{)}=1 on which the proof rests fails, because differentiating ∫W⁢f⁢dx\int Wf\,\mathrm{d}x with respect to θ\theta picks up a term from the moving boundary of the integral. The value θ2/n\theta^{2}/n is therefore not a bound at all for this model, and estimators based on the maximum are super-efficient in the sense of example 13.6.

∎

Solution to exercise 31.4.
  1. (a)
    ​

    The tray comes from supplier A with probability π\pi, and given the supplier the count is binomial, and thus by the law of total probability

    f⁢(x;π)=π⁢fA⁢(x)+(1−π)⁢fB⁢(x),x∈{0,1,…,10},f(x;\pi)=\pi f_{A}(x)+(1-\pi)f_{B}(x),\qquad x\in\{0,1,\dots,10\},

    a two-component mixture with known components. For nn independent trays the observed-data log-likelihood is

    ℓ⁢(π)\displaystyle\ell(\pi)
    =∑i=1nlog⁡(π⁢fA⁢(xi)+(1−π)⁢fB⁢(xi)),\displaystyle=\sum_{i=1}^{n}\log\bigl{(}\pi f_{A}(x_{i})+(1-\pi)f_{B}(x_{i})% \bigr{)},
    ℓ′⁢(π)\displaystyle\ell^{\prime}(\pi)
    =∑i=1nfA⁢(xi)−fB⁢(xi)π⁢fA⁢(xi)+(1−π)⁢fB⁢(xi).\displaystyle=\sum_{i=1}^{n}\frac{f_{A}(x_{i})-f_{B}(x_{i})}{\pi f_{A}(x_{i})+% (1-\pi)f_{B}(x_{i})}.

    The likelihood equation ℓ′⁢(π)=0\ell^{\prime}(\pi)=0 is, after clearing denominators, a polynomial equation of degree n−1n-1 in π\pi, and the sum inside the logarithm prevents the simplification which made the examples of lecture 5 solvable. The complete data are the pairs (Xi,Zi)(X_{i},Z_{i}) with fc⁢(x,1;π)=π⁢fA⁢(x)f_{c}(x,1;\pi)=\pi f_{A}(x) and fc⁢(x,2;π)=(1−π)⁢fB⁢(x)f_{c}(x,2;\pi)=(1-\pi)f_{B}(x), and thus

    ℓc⁢(π)=∑i=1n[𝟏{Zi=1}⁢(log⁡π+log⁡fA⁢(xi))+𝟏{Zi=2}⁢(log⁡(1−π)+log⁡fB⁢(xi))],\ell_{c}(\pi)=\sum_{i=1}^{n}\Bigl{[}\mathbf{1}_{\{Z_{i}=1\}}\bigl{(}\log\pi+% \log f_{A}(x_{i})\bigr{)}+\mathbf{1}_{\{Z_{i}=2\}}\bigl{(}\log(1-\pi)+\log f_{% B}(x_{i})\bigr{)}\Bigr{]},

    which is linear in the indicators.

  2. (b)
    ​

    By definition 19.1 the E step computes Q⁢(π|πk)=𝔼πk⁢(ℓc⁢(π)|X)Q(\pi\mskip 1.0mu|\mskip 1.0mu\pi_{k})=\mathbb{E}_{\pi_{k}}\bigl{(}\ell_{c}(% \pi)\!\mathrel{\big{|}}\!X\bigr{)}. Since ℓc\ell_{c} is linear in the indicators, and since the pairs are independent, this replaces 𝟏{Zi=1}\mathbf{1}_{\{Z_{i}=1\}} by its conditional expectation ℙπk⁢(Zi=1|Xi=xi)\mathbb{P}_{\pi_{k}}(Z_{i}=1\mskip 1.0mu|\mskip 1.0muX_{i}=x_{i}), which by Bayes’ rule, proposition A.20, is

    γi=ℙπk⁢(Zi=1)⁢ℙ⁢(Xi=xi|Zi=1)ℙπk⁢(Xi=xi)=πk⁢fA⁢(xi)πk⁢fA⁢(xi)+(1−πk)⁢fB⁢(xi),\gamma_{i}=\frac{\mathbb{P}_{\pi_{k}}(Z_{i}=1)\,\mathbb{P}(X_{i}=x_{i}\mskip 1% .0mu|\mskip 1.0muZ_{i}=1)}{\mathbb{P}_{\pi_{k}}(X_{i}=x_{i})}=\frac{\pi_{k}f_{% A}(x_{i})}{\pi_{k}f_{A}(x_{i})+(1-\pi_{k})f_{B}(x_{i})},

    and 𝟏{Zi=2}\mathbf{1}_{\{Z_{i}=2\}} by 1−γi1-\gamma_{i}. Thus

    Q⁢(π|πk)=∑i=1n[γi⁢log⁡π+(1−γi)⁢log⁡(1−π)]+terms not involving ⁢π.Q(\pi\mskip 1.0mu|\mskip 1.0mu\pi_{k})=\sum_{i=1}^{n}\Bigl{[}\gamma_{i}\log\pi% +(1-\gamma_{i})\log(1-\pi)\Bigr{]}+\text{terms not involving }\pi.
  3. (c)
    ​

    Differentiating with the γi\gamma_{i} held fixed gives

    ∂Q∂π=∑iγiπ−∑i(1−γi)1−π=∑iγi−n⁢ππ⁢(1−π),\frac{\partial Q}{\partial\pi}=\frac{\sum_{i}\gamma_{i}}{\pi}-\frac{\sum_{i}(1% -\gamma_{i})}{1-\pi}=\frac{\sum_{i}\gamma_{i}-n\pi}{\pi(1-\pi)},

    which vanishes at πk+1=1n⁢∑iγi\pi_{k+1}=\frac{1}{n}\sum_{i}\gamma_{i}, is positive to the left of this value and negative to the right. Thus πk+1\pi_{k+1} is the maximiser; equivalently, the second derivative −∑iγi/π2−∑i(1−γi)/(1−π)2-\sum_{i}\gamma_{i}/\pi^{2}-\sum_{i}(1-\gamma_{i})/(1-\pi)^{2} is negative for every π∈(0,1)\pi\in(0,1), so Q⁢(⋅|πk)Q(\,\mathord{\cdot}\,\mskip 1.0mu|\mskip 1.0mu\pi_{k}) is concave on (0,1)(0,1) and πk+1\pi_{k+1} is its global maximiser. The update is the complete-data MLE of example 5.3 with the unobserved indicators replaced by responsibilities.

  4. (d)
    ​

    With π0=1/2\pi_{0}=1/2 the weights cancel and γi=fA⁢(xi)/(fA⁢(xi)+fB⁢(xi))\gamma_{i}=f_{A}(x_{i})/\bigl{(}f_{A}(x_{i})+f_{B}(x_{i})\bigr{)}. From the table we get, to four decimal places,

    γ=(0.9952, 0.0029, 0.9995, 0.0003, 0.9572)for ⁢x=(2,7,1,8,3),\gamma=(0.9952,\ 0.0029,\ 0.9995,\ 0.0003,\ 0.9572)\qquad\text{for }x=(2,7,1,8% ,3),

    and thus ∑iγi=2.9551\sum_{i}\gamma_{i}=2.9551 and π1=2.9551/5=0.5910\pi_{1}=2.9551/5=0.5910. The responsibilities are all close to 0 or 11. Since the two component distributions barely overlap, each count identifies its supplier almost with certainty, and the trays with 11, 22 and 33 germinated seeds are assigned to supplier A, the trays with 77 and 88 to supplier B. The tray with 33 seeds is the least certain, at 0.9570.957. Further iterations move πk\pi_{k} only slightly, to the fixed point 0.5940.594; the observed-data log-likelihood rises from −10.30-10.30 at π0\pi_{0} to −10.22-10.22 at π1\pi_{1}, as theorem 19.2 requires.

  5. (e)
    ​

    With pAp_{A} unknown, log⁡fA⁢(xi)=xi⁢log⁡pA+(m−xi)⁢log⁡(1−pA)+const\log f_{A}(x_{i})=x_{i}\log p_{A}+(m-x_{i})\log(1-p_{A})+\text{const} can no longer be dropped from QQ, which now contains the term

    ∑iγi⁢(xi⁢log⁡pA+(m−xi)⁢log⁡(1−pA)).\sum_{i}\gamma_{i}\bigl{(}x_{i}\log p_{A}+(m-x_{i})\log(1-p_{A})\bigr{)}.

    Setting the derivative of this term with respect to pAp_{A} to zero gives

    ∑iγi⁢xipA−∑iγi⁢(m−xi)1−pA=0,i.e.pA,k+1=∑iγi⁢xim⁢∑iγi,\frac{\sum_{i}\gamma_{i}x_{i}}{p_{A}}-\frac{\sum_{i}\gamma_{i}(m-x_{i})}{1-p_{% A}}=0,\qquad\text{{i.e.}}\qquad p_{A,k+1}=\frac{\sum_{i}\gamma_{i}x_{i}}{m\sum% _{i}\gamma_{i}},

    the responsibility-weighted proportion of germinated seeds among the trays attributed to supplier A, with the responsibilities now computed from (πk,pA,k)(\pi_{k},p_{A,k}). For the data of part (d) and the responsibilities computed there this gives pA,1=5.884/29.55=0.199p_{A,1}=5.884/29.55=0.199. The ascent property, theorem 19.2, states that the observed-data log-likelihood never decreases along the EM sequence, ℓ⁢(πk+1,pA,k+1)≥ℓ⁢(πk,pA,k)\ell(\pi_{k+1},p_{A,k+1})\geq\ell(\pi_{k},p_{A,k}) for all kk. It guarantees that the sequence of log-likelihood values converges, and by proposition 19.3 a point at which the iteration comes to rest is a stationary point of ℓ\ell; it does not guarantee that this point is the global maximum, since a poor starting value can lead to a local maximum, and it says nothing about the speed of convergence.

∎