Lecture 31 The Exponential Model: From Start to Finish
In this final lecture we will consider the exponential distribution with unknown rate. We will follow the approach a statistician would take when considering this model: we will consider the model and the data, how to estimate the parameter, how well the estimate is performing, what intervals and tests we can use, and how things change when we consider Bayesian inference. The individual steps of this approach are based on results we have seen in previous lectures, and the aim of this lecture is to see how these results can be combined to consider an example. We choose the exponential distribution because this is the running example of the module, and because most of the individual steps in the analysis should be already known to you from lectures and exercises.
31.1 The Model and the Data
Assume that the lifetime of a component of a machine is exponentially distributed with rate . Then the lifetimes of components are i.i.d. with density for . The mean lifetime is , the variance is (table A.2). In example 7.4 we have seen that this is a one-parameter exponential family with natural statistic and natural parameter . Thus, by lemma 7.5, the joint density of the sample is
for . Since this density depends on the data only through , we can use the factorisation theorem 8.2 to conclude that is sufficient for . Since the family has full rank in the sense of definition 7.10 (the natural parameter space is the open half-line), we can use proposition 10.5 to conclude that is also complete. Thus, from lectures 7 to 10 we know that all our results can be based on alone, and that any unbiased function of will be the best unbiased estimator for the corresponding quantity.
As data we have observed lifetimes (in years) of the components, and the sum of these values is , i.e. we have . The manufacturer claims that the mean lifetime is two years, i.e. , and we want to test whether the components of the machine last longer than the manufacturer claims.
31.2 Point Estimation
The log-likelihood is , with derivative and second derivative , so the likelihood equation has the unique root and this root is the maximum: is the MLE for (example 5.2). For our data we find . By the invariance of the MLE, theorem 5.6, the MLE for the mean lifetime is years.
The MLE for the rate is biased: exercise 2.3 shows , so that overestimates by a factor of . The corrected estimator is unbiased. Since is an unbiased function of the complete sufficient statistic , the Lehmann–Scheffe theorem 10.4 shows that this is the UMVUE for . The variance of the UMVUE can be found as (exercise 5.3). For the data considered here, we find .
What is the optimal performance of an estimator? From example 11.8 we know that the Fisher information of one observation is and thus we have . Since the model is regular (exercise 16.1), the Cramer–Rao inequality (theorem 13.2) gives the bound for the variance of any unbiased estimator. The variance of the UMVUE is , which is larger than the bound and by proposition 13.5 no unbiased estimator can attain the bound, since the bound is only attained for an affine function of the natural statistic and no such function is unbiased for the rate. The efficiency of tends to as . The situation for the mean is different: is unbiased with variance . This variance equals the bound for the mean (exercise 11.5) and thus is efficient.
31.3 Interval Estimation
Proposition 25.7 turns the standard error into the approximate confidence interval . For the data considered here, this confidence interval is , i.e. . The interval is symmetric about the estimate, and the probability of the interval covering the correct value is only approximately , since the interval is based on the normal limit at .
For the exponential distribution we can do better, because an exact pivot is available. In exercise 13.1, written there in terms of the mean , we have seen that for every value of , so is a pivot in the sense of lecture 25. Solving for gives the exact interval
with coverage exactly for every . With and we find the interval . The exact interval is shifted to the right relative to the Wald interval and is not symmetric about : it makes use of the fact that the sampling distribution of is skewed, which the normal approximation does not. Both intervals are confidence intervals in the sense of definition 25.1: the statement is about the procedure, and for the interval in front of us the parameter is either inside the interval or it is not.
31.4 Testing
We now test the manufacturer’s claim. We take and start, as in lecture 22, with a simple alternative where (a longer mean lifetime). The likelihood ratio is
which is an increasing function of , since . By the Neyman–Pearson lemma, theorem 22.2, the most powerful test at level rejects if , where is chosen such that . Under the pivot gives , and thus . For we find . The critical region does not depend on and thus, by proposition 22.4, the same test is most powerful against all alternatives . Since the observed value is less than the critical value, we do not reject at the level: the data do not show that the components have a longer lifetime than claimed. The -value from definition 20.8 is .
The power function of definition 20.4 can be found by evaluating the same pivot at the alternative: , so for example . If the true mean lifetime of the components is four years, then a sample of ten components has probability of detecting this.
For the two-sided alternative there is no uniformly most powerful test, by the argument of proposition 22.5: the two one-sided tests reject on opposite tails of . We use the generalised likelihood-ratio test of definition 23.1 instead. The statistic is
and by Wilks’ theorem 23.3 it is approximately -distributed under , since one parameter is fixed. For our data, , well below , so again is not rejected. The exercises at the end of this lecture verify the formula for .
31.5 The Bayesian Answer
Since we can assume to be random with prior density , and since the likelihood function is as a function of , a prior of the same shape as the likelihood, namely the Gamma density , is conjugate. This is a special case of theorem 26.2: the posterior distribution is given by
which is again the Gamma distribution. The situation in the example fits the result of proposition 26.4: the update process increases by and by , so that the effect of the prior is as if we had observed additional values with total lifetime . In the example of the manufacturer’s claim about the lifetime of the components, we can assume that the claim is equivalent to having observed two more prior observations with mean lifetime two years, i.e. and . The prior mean is then . The posterior distribution is then Gamma and, using theorem 26.7, the Bayes estimator for , using squared error loss, is given by
This is a weighted average of the prior mean and the MLE, where the weight of the data increases as increases.
The credible interval of definition 28.1 comes from the same chi-squared fact as before: if then , so the equal-tailed credible interval is . For the manufacturer’s claim, the posterior probability of long lifetimes is , up from the prior probability . Unlike the -value, this is a probability about the parameter, computed after the data have been seen.
The Jeffreys prior from definition 26.8 is . This is the improper limit of the Gamma prior. Using this prior, the posterior distribution is Gamma. The credible interval is . This is the same as the exact confidence interval introduced in section 31.3. Also, . This is the -value of the one-sided test. While the numbers coincide, the statements are different: the confidence interval covers the fixed in of repetitions, whereas the credible interval covers the random with probability given the data; and the -value is a probability about the data under , whereas is a probability about the hypothesis. This illustrates the comparison from section 28.3, for a concrete model, and it is special that the numbers coincide for the Jeffreys prior and one-sided hypotheses.
31.6 Looking Back and Looking Forward
Every question about the exponential model we have been able to answer so far has used a small number of different tools: the exponential-family structure of the model allowed us to find a complete sufficient statistic, the likelihood function allowed us to find the estimator, the Fisher information allowed us to find the standard error and the bound, a pivot or the normal limit allowed us to find the interval, the Neyman–Pearson lemma and Wilks’ theorem allowed us to find tests, and Bayes’ theorem allowed us to find the posterior distribution. The same list of tools applies for any model, and it is this list which you should use when you are given a model not considered in these notes: Write down the likelihood function, find the sufficient statistic, solve the likelihood equation and check that the root is a maximum, compute the information, and use the results to compute the standard error, the interval, and the tests. The exact pivot is a luxury of the exponential distribution; in general you will only have the large-sample results from lectures 16, 17 and 23, and some care is needed when these results are used.
This is also the material we will study after this module. The standard errors computed by all regression packages are the inverse Fisher information from lecture 14, the tests which compare fitted models from two different analyses are the likelihood-ratio tests from lecture 23, mixture models and hidden Markov models are fitted using the EM algorithm from lecture 19, and Bayesian computation, from the simple conjugate updating to the simulation methods of later modules, is based on the posterior distribution from lecture 26. The models under consideration here will be much larger, but the questions we ask of these models will be similar to the ones we considered in this module.
-
•
For the exponential distribution with rate , the sum is complete and sufficient, the MLE is , the UMVUE is , and no unbiased estimator can achieve the Cramer–Rao bound , although is efficient for the mean.
-
•
The standard error can be used to compute the Wald interval. The pivot can be used to compute an exact interval, an exact one-sided test with known power function, and the -value.
-
•
A two-sided test can be based on and Wilks’ theorem.
-
•
For the Gamma prior, the posterior distribution is Gamma, the posterior mean is a weighted average of the prior mean and the MLE, and the Jeffreys prior gives the exact interval and -value numerically, but with a different meaning.
-
•
The same short list of tools can be used to answer all questions about all regular models.
Let be i.i.d. with exponential rate and let be the MLE.
-
1.
Show that the generalised likelihood-ratio statistic for against is given by , and that depends on the data only through , where .
-
2.
Show that is zero at and positive elsewhere, i.e. that with equality if and only if .
-
3.
For and , determine the value of for and for and state the conclusion of the test at level for both cases.
Let be i.i.d. exponential with rate and let .
-
1.
Show that the Jeffreys prior for is . Show that this prior is improper.
-
2.
Show that the posterior distribution under the Jeffreys prior is proper for every sample with . Identify the posterior distribution as a Gamma distribution.
-
3.
Using proposition 26.9 or otherwise, find the Jeffreys prior for the mean . Show that this prior is the image of under the map .
-
4.
Show that the equal-tailed credible interval for , under the Jeffreys prior, coincides with the exact confidence interval from section 31.3, but in one sentence explain why the two intervals make different statements.
Let with be i.i.d. uniform on , where is unknown, and let and . Without proof, we state that, given , the remaining observations are i.i.d. uniform on and .
-
(a)
Find the density of and show that .
-
(b)
Show that is an unbiased estimator for . Determine and show that is not a consistent estimator for .
-
(c)
Show that is sufficient for .
-
(d)
Determine and show that . Which theorem guarantees, without computation, that is a statistic, unbiased for and with a variance of at most that of ?
-
(e)
Explain why is the UMVUE for . The Cramer–Rao formula, for this model, gives the value , which is larger than the variance of . Explain why there is no contradiction.
Seed trays are bought from two suppliers. Each tray contains seeds, and each seed has probability of germinating if the tray comes from supplier A, and probability if the tray comes from supplier B, independently of the other seeds. Both probabilities are known. It is unknown what proportion of trays come from supplier A, but the supplier of individual trays is not recorded. For tray , the number of germinated seeds is observed, and we write if the tray came from supplier A and otherwise. Let and be the probability weights of the and distributions.
-
(a)
Write down the probability weights of the number of germinated seeds and the observed-data log-likelihood for trays. Explain why the likelihood equation cannot be solved analytically and write down the complete-data log-likelihood .
-
(b)
Derive the E step of the EM algorithm, showing that can be found from by replacing the indicator by the responsibility
-
(c)
Derive the M step, , and show that this is a maximum of .
-
(d)
For five trays, the number of germinated seeds was , , , and . The probability weights of the two distributions are given by
Given , perform one iteration of the EM algorithm and comment on your responsibilities.
-
(e)
Assume now that is unknown, and is still known. Derive the M step for . State the ascent property of the EM algorithm, and explain what the property does and does not guarantee for the sequence .