Lecture 31 The Exponential Model: From Start to Finish

In this final lecture we will consider the exponential distribution with unknown rate. We will follow the approach a statistician would take when considering this model: we will consider the model and the data, how to estimate the parameter, how well the estimate is performing, what intervals and tests we can use, and how things change when we consider Bayesian inference. The individual steps of this approach are based on results we have seen in previous lectures, and the aim of this lecture is to see how these results can be combined to consider an example. We choose the exponential distribution because this is the running example of the module, and because most of the individual steps in the analysis should be already known to you from lectures and exercises.

31.1 The Model and the Data

Assume that the lifetime of a component of a machine is exponentially distributed with rate θ>0\theta>0. Then the lifetimes X1,…,XnX_{1},\dots,X_{n} of nn components are i.i.d. with density f⁢(x;θ)=θ⁢e−θ⁢xf(x;\theta)=\theta e^{-\theta x} for x≥0x\geq 0. The mean lifetime is 𝔼θ⁢(Xi)=1/θ\mathbb{E}_{\theta}(X_{i})=1/\theta, the variance is 1/θ21/\theta^{2} (table A.2). In example 7.4 we have seen that this is a one-parameter exponential family with natural statistic T⁢(x)=xT(x)=x and natural parameter η⁢(θ)=−θ\eta(\theta)=-\theta. Thus, by lemma 7.5, the joint density of the sample is

f⁢(x1,…,xn;θ)=θn⁢exp⁡(−θ⁢∑i=1nxi)f(x_{1},\dots,x_{n};\theta)=\theta^{n}\exp\Bigl{(}-\theta\sum_{i=1}^{n}x_{i}% \Bigr{)}

for x1,…,xn≥0x_{1},\dots,x_{n}\geq 0. Since this density depends on the data only through T=∑i=1nXiT=\sum_{i=1}^{n}X_{i}, we can use the factorisation theorem 8.2 to conclude that TT is sufficient for θ\theta. Since the family has full rank in the sense of definition 7.10 (the natural parameter space is the open half-line), we can use proposition 10.5 to conclude that TT is also complete. Thus, from lectures 7 to 10 we know that all our results can be based on TT alone, and that any unbiased function of TT will be the best unbiased estimator for the corresponding quantity.

As data we have observed n=10n=10 lifetimes (in years) of the components, and the sum of these values is ∑ixi=25.0\sum_{i}x_{i}=25.0, i.e. we have x¯=2.5\bar{x}=2.5. The manufacturer claims that the mean lifetime is two years, i.e. θ0=0.5\theta_{0}=0.5, and we want to test whether the components of the machine last longer than the manufacturer claims.

31.2 Point Estimation

The log-likelihood is ℓ⁢(θ)=n⁢log⁡θ−θ⁢T\ell(\theta)=n\log\theta-\theta T, with derivative ℓ′⁢(θ)=n/θ−T\ell^{\prime}(\theta)=n/\theta-T and second derivative −n/θ2<0-n/\theta^{2}<0, so the likelihood equation has the unique root θ^=n/T=1/X¯\hat{\theta}=n/T=1/\bar{X} and this root is the maximum: θ^\hat{\theta} is the MLE for θ\theta (example 5.2). For our data we find θ^=10/25=0.4\hat{\theta}=10/25=0.4. By the invariance of the MLE, theorem 5.6, the MLE for the mean lifetime 1/θ1/\theta is X¯=2.5\bar{X}=2.5 years.

The MLE for the rate is biased: exercise 2.3 shows 𝔼θ⁢(θ^)=n⁢θ/(n−1)\mathbb{E}_{\theta}(\hat{\theta})=n\theta/(n-1), so that θ^\hat{\theta} overestimates θ\theta by a factor of n/(n−1)n/(n-1). The corrected estimator θ~=(n−1)/T\tilde{\theta}=(n-1)/T is unbiased. Since θ~\tilde{\theta} is an unbiased function of the complete sufficient statistic TT, the Lehmann–Scheffe theorem 10.4 shows that this is the UMVUE for θ\theta. The variance of the UMVUE can be found as θ2/(n−2)\theta^{2}/(n-2) (exercise 5.3). For the data considered here, we find θ~=9/25=0.36\tilde{\theta}=9/25=0.36.

What is the optimal performance of an estimator? From example 11.8 we know that the Fisher information of one observation is ℐ︀⁢(θ)=1/θ2\mathcal{I}(\theta)=1/\theta^{2} and thus we have ℐ︀n⁢(θ)=n/θ2\mathcal{I}_{n}(\theta)=n/\theta^{2}. Since the model is regular (exercise 16.1), the Cramer–Rao inequality (theorem 13.2) gives the bound θ2/n\theta^{2}/n for the variance of any unbiased estimator. The variance of the UMVUE is θ2/(n−2)\theta^{2}/(n-2), which is larger than the bound and by proposition 13.5 no unbiased estimator can attain the bound, since the bound is only attained for an affine function of the natural statistic TT and no such function is unbiased for the rate. The efficiency (n−2)/n(n-2)/n of θ~\tilde{\theta} tends to 11 as n→∞n\to\infty. The situation for the mean 1/θ1/\theta is different: X¯\bar{X} is unbiased with variance 1/(n⁢θ2)1/(n\theta^{2}). This variance equals the bound 1/ℐ︀n⁢(μ)1/\mathcal{I}_{n}(\mu) for the mean (exercise 11.5) and thus X¯\bar{X} is efficient.

For large samples, theorems 16.3 and 17.1 apply, because the regularity conditions are satisfied: θ^n\hat{\theta}_{n} is consistent and n⁢(θ^n−θ)→dN⁢(0,θ2)\sqrt{n}(\hat{\theta}_{n}-\theta)\xrightarrow{\ \mathrm{d}\ }N(0,\theta^{2}). From corollary 17.3 we know that the standard error is

se(θ^)=1ℐ︀n⁢(θ^)=θ^n.\mathop{\mathrm{se}}\nolimits(\hat{\theta})=\frac{1}{\sqrt{\mathcal{I}_{n}(% \hat{\theta})}}=\frac{\hat{\theta}}{\sqrt{n}}.

For the data at hand we get se(θ^)=0.4/10=0.126\mathop{\mathrm{se}}\nolimits(\hat{\theta})=0.4/\sqrt{10}=0.126. This is the number we would expect to see next to the estimate.

31.3 Interval Estimation

Proposition 25.7 turns the standard error into the approximate 95%95\% confidence interval θ^±1.96⁢se(θ^)\hat{\theta}\pm 1.96\,\mathop{\mathrm{se}}\nolimits(\hat{\theta}). For the data considered here, this confidence interval is 0.4±0.2480.4\pm 0.248, i.e. [0.152,0.648][0.152,0.648]. The interval is symmetric about the estimate, and the probability of the interval covering the correct value is only approximately 95%95\%, since the interval is based on the normal limit at n=10n=10.

For the exponential distribution we can do better, because an exact pivot is available. In exercise 13.1, written there in terms of the mean 1/θ1/\theta, we have seen that 2⁢θ⁢T∼χ2⁢n22\theta T\sim\chi^{2}_{2n} for every value of θ\theta, so Q=2⁢θ⁢TQ=2\theta T is a pivot in the sense of lecture 25. Solving χ2⁢n,α/22≤2⁢θ⁢T≤χ2⁢n,1−α/22\chi^{2}_{2n,\alpha/2}\leq 2\theta T\leq\chi^{2}_{2n,1-\alpha/2} for θ\theta gives the exact interval

[χ2⁢n,α/222⁢T,χ2⁢n,1−α/222⁢T]\Bigl{[}\frac{\chi^{2}_{2n,\alpha/2}}{2T},\ \frac{\chi^{2}_{2n,1-\alpha/2}}{2T% }\Bigr{]}

with coverage exactly 1−α1-\alpha for every nn. With χ20,0.0252=9.591\chi^{2}_{20,0.025}=9.591 and χ20,0.9752=34.170\chi^{2}_{20,0.975}=34.170 we find the interval [9.591/50,34.170/50]=[0.192,0.683][9.591/50,34.170/50]=[0.192,0.683]. The exact interval is shifted to the right relative to the Wald interval and is not symmetric about 0.40.4: it makes use of the fact that the sampling distribution of θ^\hat{\theta} is skewed, which the normal approximation does not. Both intervals are 95%95\% confidence intervals in the sense of definition 25.1: the statement is about the procedure, and for the interval in front of us the parameter is either inside the interval or it is not.

31.4 Testing

We now test the manufacturer’s claim. We take H0:θ=θ0=0.5H_{0}\colon\theta=\theta_{0}=0.5 and start, as in lecture 22, with a simple alternative H1:θ=θ1H_{1}\colon\theta=\theta_{1} where θ1<θ0\theta_{1}<\theta_{0} (a longer mean lifetime). The likelihood ratio is

f⁢(x;θ1)f⁢(x;θ0)=(θ1θ0)n⁢exp⁡((θ0−θ1)⁢T),\frac{f(x;\theta_{1})}{f(x;\theta_{0})}=\Bigl{(}\frac{\theta_{1}}{\theta_{0}}% \Bigr{)}^{n}\exp\bigl{(}(\theta_{0}-\theta_{1})T\bigr{)},

which is an increasing function of TT, since θ0−θ1>0\theta_{0}-\theta_{1}>0. By the Neyman–Pearson lemma, theorem 22.2, the most powerful test at level α\alpha rejects H0H_{0} if T>cT>c, where cc is chosen such that ℙθ0⁢(T>c)=α\mathbb{P}_{\theta_{0}}(T>c)=\alpha. Under H0H_{0} the pivot gives 2⁢θ0⁢T∼χ2⁢n22\theta_{0}T\sim\chi^{2}_{2n}, and thus c=χ2⁢n,1−α2/(2⁢θ0)c=\chi^{2}_{2n,1-\alpha}/(2\theta_{0}). For α=0.05\alpha=0.05 we find c=31.410/1=31.41c=31.410/1=31.41. The critical region {T>31.41}\{T>31.41\} does not depend on θ1\theta_{1} and thus, by proposition 22.4, the same test is most powerful against all alternatives θ<θ0\theta<\theta_{0}. Since the observed value T=25.0T=25.0 is less than the critical value, we do not reject H0H_{0} at the 5%5\% level: the data do not show that the components have a longer lifetime than claimed. The pp-value from definition 20.8 is ℙθ0⁢(T≥25.0)=ℙ⁢(χ202≥25.0)=0.20\mathbb{P}_{\theta_{0}}(T\geq 25.0)=\mathbb{P}(\chi^{2}_{20}\geq 25.0)=0.20.

The power function of definition 20.4 can be found by evaluating the same pivot at the alternative: β⁢(θ)=ℙθ⁢(T>c)=ℙ⁢(χ2⁢n2>2⁢θ⁢c)\beta(\theta)=\mathbb{P}_{\theta}(T>c)=\mathbb{P}(\chi^{2}_{2n}>2\theta c), so for example β⁢(0.25)=ℙ⁢(χ202>15.71)=0.73\beta(0.25)=\mathbb{P}(\chi^{2}_{20}>15.71)=0.73. If the true mean lifetime of the components is four years, then a sample of ten components has probability 0.730.73 of detecting this.

For the two-sided alternative θ≠θ0\theta\neq\theta_{0} there is no uniformly most powerful test, by the argument of proposition 22.5: the two one-sided tests reject on opposite tails of TT. We use the generalised likelihood-ratio test of definition 23.1 instead. The statistic is

W=2⁢(ℓ⁢(θ^)−ℓ⁢(θ0))=2⁢n⁢(log⁡θ^θ0−1+θ0θ^),W=2\bigl{(}\ell(\hat{\theta})-\ell(\theta_{0})\bigr{)}=2n\Bigl{(}\log\frac{% \hat{\theta}}{\theta_{0}}-1+\frac{\theta_{0}}{\hat{\theta}}\Bigr{)},

and by Wilks’ theorem 23.3 it is approximately χ12\chi^{2}_{1}-distributed under H0H_{0}, since one parameter is fixed. For our data, W=20⁢(log⁡0.8−1+1.25)=0.54W=20\bigl{(}\log 0.8-1+1.25\bigr{)}=0.54, well below χ1,0.952=3.84\chi^{2}_{1,0.95}=3.84, so again H0H_{0} is not rejected. The exercises at the end of this lecture verify the formula for WW.

31.5 The Bayesian Answer

Since we can assume θ\theta to be random with prior density π⁢(θ)\pi(\theta), and since the likelihood function is θn⁢e−θ⁢T\theta^{n}e^{-\theta T} as a function of θ\theta, a prior of the same shape as the likelihood, namely the Gamma(a,b)(a,b) density π⁢(θ)∝θa−1⁢e−b⁢θ\pi(\theta)\propto\theta^{a-1}e^{-b\theta}, is conjugate. This is a special case of theorem 26.2: the posterior distribution is given by

π⁢(θ|x)∝θa−1⁢e−b⁢θ⁢θn⁢e−θ⁢T=θa+n−1⁢e−(b+T)⁢θ,\pi(\theta\mskip 1.0mu|\mskip 1.0mux)\propto\theta^{a-1}e^{-b\theta}\,\theta^{% n}e^{-\theta T}=\theta^{a+n-1}e^{-(b+T)\theta},

which is again the Gamma(a+n,b+T)(a+n,b+T) distribution. The situation in the example fits the result of proposition 26.4: the update process increases aa by nn and bb by TT, so that the effect of the prior is as if we had observed aa additional values with total lifetime bb. In the example of the manufacturer’s claim about the lifetime of the components, we can assume that the claim is equivalent to having observed two more prior observations with mean lifetime two years, i.e. a=2a=2 and b=4b=4. The prior mean is then a/b=0.5a/b=0.5. The posterior distribution is then Gamma(12,29)(12,29) and, using theorem 26.7, the Bayes estimator for θ\theta, using squared error loss, is given by

𝔼⁢(θ|x)=a+nb+T=bb+T⋅ab+Tb+T⋅nT=0.138⋅0.5+0.862⋅0.4=0.414.\mathbb{E}(\theta\mskip 1.0mu|\mskip 1.0mux)=\frac{a+n}{b+T}=\frac{b}{b+T}% \cdot\frac{a}{b}+\frac{T}{b+T}\cdot\frac{n}{T}=0.138\cdot 0.5+0.862\cdot 0.4=0% .414.

This is a weighted average of the prior mean and the MLE, where the weight of the data increases as TT increases.

The credible interval of definition 28.1 comes from the same chi-squared fact as before: if θ∼Gamma⁢(a+n,b+T)\theta\sim\text{Gamma}(a+n,b+T) then 2⁢(b+T)⁢θ∼χ2⁢(a+n)22(b+T)\theta\sim\chi^{2}_{2(a+n)}, so the equal-tailed 95%95\% credible interval is [χ24,0.0252,χ24,0.9752]/(2⋅29)=[0.214,0.679][\chi^{2}_{24,0.025},\chi^{2}_{24,0.975}]/(2\cdot 29)=[0.214,0.679]. For the manufacturer’s claim, the posterior probability of long lifetimes is ℙ⁢(θ⁢<0.5|⁢x)=ℙ⁢(χ242<29)=0.78\mathbb{P}(\theta<0.5\mskip 1.0mu|\mskip 1.0mux)=\mathbb{P}(\chi^{2}_{24}<29)=% 0.78, up from the prior probability 0.590.59. Unlike the pp-value, this is a probability about the parameter, computed after the data have been seen.

The Jeffreys prior from definition 26.8 is πJ⁢(θ)∝ℐ︀⁢(θ)=1/θ\pi_{J}(\theta)\propto\sqrt{\mathcal{I}(\theta)}=1/\theta. This is the improper limit a,b→0a,b\to 0 of the Gamma prior. Using this prior, the posterior distribution is Gamma(n,T)(n,T). The credible interval is [χ2⁢n,α/22,χ2⁢n,1−α/22]/(2⁢T)[\chi^{2}_{2n,\alpha/2},\chi^{2}_{2n,1-\alpha/2}]/(2T). This is the same as the exact confidence interval introduced in section 31.3. Also, ℙ⁢(θ≥θ0|x)=ℙ⁢(χ2⁢n2≥2⁢θ0⁢T)\mathbb{P}(\theta\geq\theta_{0}\mskip 1.0mu|\mskip 1.0mux)=\mathbb{P}(\chi^{2}% _{2n}\geq 2\theta_{0}T). This is the pp-value of the one-sided test. While the numbers coincide, the statements are different: the confidence interval covers the fixed θ\theta in 95%95\% of repetitions, whereas the credible interval covers the random θ\theta with probability 0.950.95 given the data; and the pp-value is a probability about the data under H0H_{0}, whereas ℙ⁢(θ≥θ0|x)\mathbb{P}(\theta\geq\theta_{0}\mskip 1.0mu|\mskip 1.0mux) is a probability about the hypothesis. This illustrates the comparison from section 28.3, for a concrete model, and it is special that the numbers coincide for the Jeffreys prior and one-sided hypotheses.

31.6 Looking Back and Looking Forward

Every question about the exponential model we have been able to answer so far has used a small number of different tools: the exponential-family structure of the model allowed us to find a complete sufficient statistic, the likelihood function allowed us to find the estimator, the Fisher information allowed us to find the standard error and the bound, a pivot or the normal limit allowed us to find the interval, the Neyman–Pearson lemma and Wilks’ theorem allowed us to find tests, and Bayes’ theorem allowed us to find the posterior distribution. The same list of tools applies for any model, and it is this list which you should use when you are given a model not considered in these notes: Write down the likelihood function, find the sufficient statistic, solve the likelihood equation and check that the root is a maximum, compute the information, and use the results to compute the standard error, the interval, and the tests. The exact pivot is a luxury of the exponential distribution; in general you will only have the large-sample results from lectures 16, 17 and 23, and some care is needed when these results are used.

This is also the material we will study after this module. The standard errors computed by all regression packages are the inverse Fisher information from lecture 14, the tests which compare fitted models from two different analyses are the likelihood-ratio tests from lecture 23, mixture models and hidden Markov models are fitted using the EM algorithm from lecture 19, and Bayesian computation, from the simple conjugate updating to the simulation methods of later modules, is based on the posterior distribution from lecture 26. The models under consideration here will be much larger, but the questions we ask of these models will be similar to the ones we considered in this module.

Summary.
  • •
    ​

    For the exponential distribution with rate θ\theta, the sum TT is complete and sufficient, the MLE is 1/X¯1/\bar{X}, the UMVUE is (n−1)/T(n-1)/T, and no unbiased estimator can achieve the Cramer–Rao bound θ2/n\theta^{2}/n, although X¯\bar{X} is efficient for the mean.

  • •
    ​

    The standard error θ^/n\hat{\theta}/\sqrt{n} can be used to compute the Wald interval. The pivot 2⁢θ⁢T∼χ2⁢n22\theta T\sim\chi^{2}_{2n} can be used to compute an exact interval, an exact one-sided test with known power function, and the pp-value.

  • •
    ​

    A two-sided test can be based on W=2⁢n⁢(log⁡(θ^/θ0)−1+θ0/θ^)W=2n(\log(\hat{\theta}/\theta_{0})-1+\theta_{0}/\hat{\theta}) and Wilks’ theorem.

  • •
    ​

    For the Gamma(a,b)(a,b) prior, the posterior distribution is Gamma(a+n,b+T)(a+n,b+T), the posterior mean is a weighted average of the prior mean and the MLE, and the Jeffreys prior 1/θ1/\theta gives the exact interval and pp-value numerically, but with a different meaning.

  • •
    ​

    The same short list of tools can be used to answer all questions about all regular models.

Exercise 31.1.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with exponential rate θ>0\theta>0 and let θ^=1/X¯\hat{\theta}=1/\bar{X} be the MLE.

  1. 1.
    ​

    Show that the generalised likelihood-ratio statistic for H0:θ=θ0H_{0}\colon\theta=\theta_{0} against H1:θ≠θ0H_{1}\colon\theta\neq\theta_{0} is given by W=2⁢n⁢(log⁡(θ^/θ0)−1+θ0/θ^)W=2n\bigl{(}\log(\hat{\theta}/\theta_{0})-1+\theta_{0}/\hat{\theta}\bigr{)}, and that WW depends on the data only through u=θ0⁢X¯u=\theta_{0}\bar{X}, where W=2⁢n⁢(u−1−log⁡u)W=2n(u-1-\log u).

  2. 2.
    ​

    Show that u↦u−1−log⁡uu\mapsto u-1-\log u is zero at u=1u=1 and positive elsewhere, i.e. that W≥0W\geq 0 with equality if and only if θ^=θ0\hat{\theta}=\theta_{0}.

  3. 3.
    ​

    For n=10n=10 and x¯=2.5\bar{x}=2.5, determine the value of WW for θ0=0.5\theta_{0}=0.5 and for θ0=1\theta_{0}=1 and state the conclusion of the test at level 0.050.05 for both cases.

Exercise 31.2.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. exponential with rate θ>0\theta>0 and let T=∑i=1nXiT=\sum_{i=1}^{n}X_{i}.

  1. 1.
    ​

    Show that the Jeffreys prior for θ\theta is πJ⁢(θ)∝1/θ\pi_{J}(\theta)\propto 1/\theta. Show that this prior is improper.

  2. 2.
    ​

    Show that the posterior distribution under the Jeffreys prior is proper for every sample with T>0T>0. Identify the posterior distribution as a Gamma distribution.

  3. 3.
    ​

    Using proposition 26.9 or otherwise, find the Jeffreys prior for the mean μ=1/θ\mu=1/\theta. Show that this prior is the image of πJ\pi_{J} under the map θ↦1/θ\theta\mapsto 1/\theta.

  4. 4.
    ​

    Show that the equal-tailed 95%95\% credible interval for θ\theta, under the Jeffreys prior, coincides with the exact confidence interval from section 31.3, but in one sentence explain why the two intervals make different statements.

Exercise 31.3.

Let X1,…,XnX_{1},\dots,X_{n} with n≥2n\geq 2 be i.i.d. uniform on (0,θ)(0,\theta), where θ>0\theta>0 is unknown, and let X(1)=mini⁡XiX_{(1)}=\min_{i}X_{i} and X(n)=maxi⁡XiX_{(n)}=\max_{i}X_{i}. Without proof, we state that, given X(n)=mX_{(n)}=m, the remaining n−1n-1 observations are i.i.d. uniform on (0,m)(0,m) and Varθ((n+1)⁢X(n)/n)=θ2/(n⁢(n+2))\mathop{\mathrm{Var}}\nolimits_{\theta}\bigl{(}(n+1)X_{(n)}/n\bigr{)}=\theta^{% 2}/\bigl{(}n(n+2)\bigr{)}.

  1. (a)
    ​

    Find the density of X(1)X_{(1)} and show that 𝔼θ⁢(X(1))=θ/(n+1)\mathbb{E}_{\theta}(X_{(1)})=\theta/(n+1).

  2. (b)
    ​

    Show that U=(n+1)⁢X(1)U=(n+1)X_{(1)} is an unbiased estimator for θ\theta. Determine Varθ(U)\mathop{\mathrm{Var}}\nolimits_{\theta}(U) and show that UU is not a consistent estimator for θ\theta.

  3. (c)
    ​

    Show that X(n)X_{(n)} is sufficient for θ\theta.

  4. (d)
    ​

    Determine V=𝔼θ⁢(U|X(n))V=\mathbb{E}_{\theta}(U\mskip 1.0mu|\mskip 1.0muX_{(n)}) and show that V=(n+1)⁢X(n)/nV=(n+1)X_{(n)}/n. Which theorem guarantees, without computation, that VV is a statistic, unbiased for θ\theta and with a variance of at most that of UU?

  5. (e)
    ​

    Explain why VV is the UMVUE for θ\theta. The Cramer–Rao formula, for this model, gives the value θ2/n\theta^{2}/n, which is larger than the variance of VV. Explain why there is no contradiction.

Exercise 31.4.

Seed trays are bought from two suppliers. Each tray contains m=10m=10 seeds, and each seed has probability pA=0.2p_{A}=0.2 of germinating if the tray comes from supplier A, and probability pB=0.7p_{B}=0.7 if the tray comes from supplier B, independently of the other seeds. Both probabilities are known. It is unknown what proportion π∈(0,1)\pi\in(0,1) of trays come from supplier A, but the supplier of individual trays is not recorded. For tray ii, the number XiX_{i} of germinated seeds is observed, and we write Zi=1Z_{i}=1 if the tray came from supplier A and Zi=2Z_{i}=2 otherwise. Let fAf_{A} and fBf_{B} be the probability weights of the Binomial⁢(10,0.2)\text{Binomial}(10,0.2) and Binomial⁢(10,0.7)\text{Binomial}(10,0.7) distributions.

  1. (a)
    ​

    Write down the probability weights of the number XiX_{i} of germinated seeds and the observed-data log-likelihood ℓ⁢(π)\ell(\pi) for nn trays. Explain why the likelihood equation cannot be solved analytically and write down the complete-data log-likelihood ℓc⁢(π)\ell_{c}(\pi).

  2. (b)
    ​

    Derive the E step of the EM algorithm, showing that Q⁢(π|πk)Q(\pi\mskip 1.0mu|\mskip 1.0mu\pi_{k}) can be found from ℓc⁢(π)\ell_{c}(\pi) by replacing the indicator 𝟏{Zi=1}\mathbf{1}_{\{Z_{i}=1\}} by the responsibility

    γi=πk⁢fA⁢(xi)πk⁢fA⁢(xi)+(1−πk)⁢fB⁢(xi).\gamma_{i}=\frac{\pi_{k}f_{A}(x_{i})}{\pi_{k}f_{A}(x_{i})+(1-\pi_{k})f_{B}(x_{% i})}.
  3. (c)
    ​

    Derive the M step, πk+1=1n⁢∑i=1nγi\pi_{k+1}=\frac{1}{n}\sum_{i=1}^{n}\gamma_{i}, and show that this is a maximum of π↦Q⁢(π|πk)\pi\mapsto Q(\pi\mskip 1.0mu|\mskip 1.0mu\pi_{k}).

  4. (d)
    ​

    For five trays, the number of germinated seeds was 22, 77, 11, 88 and 33. The probability weights of the two distributions are given by

    x12378fA⁢(x)0.26840.30200.20130.0007860.0000737fB⁢(x)0.0001380.001450.009000.26680.2335\begin{array}[]{lccccc}x&1&2&3&7&8\\ f_{A}(x)&0.2684&0.3020&0.2013&0.000786&0.0000737\\ f_{B}(x)&0.000138&0.00145&0.00900&0.2668&0.2335\end{array}

    Given π0=1/2\pi_{0}=1/2, perform one iteration of the EM algorithm and comment on your responsibilities.

  5. (e)
    ​

    Assume now that pAp_{A} is unknown, and pBp_{B} is still known. Derive the M step for pAp_{A}. State the ascent property of the EM algorithm, and explain what the property does and does not guarantee for the sequence (πk,pA,k)(\pi_{k},p_{A,k}).