Preparing for the Exam
This chapter summarises what you need to do to prepare for the examination: it includes a map of the module, a summary of important results, information about what you were required to prove and what you were only required to know, advice on how to write your answers, details of common mistakes to avoid, and information on how to use the practice paper. None of this is new material, but it is gathered together here from different parts of the module for revision purposes.
The Module in One Page
The module consists of one question, split into two parts. The first part, covering lectures 1 to 19, deals with estimation: the question is what an estimator is, and how one can measure quality of an estimator. In this part of the module we also learn how estimators can be constructed using the method of moments and the maximum likelihood method, how a sufficient statistic can be used to compress the data and how conditioning on the data can be used to improve an estimator. In this part of the module we also learn about the Fisher information and how it can be used to give a bound on the variance of all unbiased estimators. We learn how the maximum likelihood estimator behaves for large sample size and how to compute the maximum likelihood estimator when some of the data are missing. Throughout this part of the module we will consider the exponential families, introduced in lecture 7. These families of distributions provide the sufficient statistic, the completeness of the sufficient statistic, the Fisher information and, in lecture 26, the conjugate priors. The second part of the module, covering lectures 20 to 29, deals with testing and interval estimation. In this part of the module we learn about the language of tests, the best test for simple hypotheses, the likelihood ratio for composite hypotheses, the duality of tests and confidence intervals, the Bayesian approach to the same questions and the standard tests for normal data. Table 31.1 shows the lectures for each week of the module and the results of each lecture and the later lectures which depend on the results of these lectures.
| Week | Lectures and main results | Used later in |
| 1 | 1: model, statistic, estimator (definitions 1.1, 1.3 and 1.4). 2: bias and mean squared error, theorem 2.7; consistency, theorem 2.10. | every later lecture; the bias corrections of lecture 5; the consistency proof of lecture 16 |
| 2 | 4: likelihood, score and method of moments (definitions 4.1, 4.4 and 4.6). 5: the MLE (definition 5.1), the likelihood equation and its check, invariance, theorem 5.6; failures, section 5.3. | lectures 11, 16, 17, 19 and 23 |
| 3 | 7: exponential families (definition 7.1), the cumulant function, theorem 7.7, full rank (definition 7.10). 8: sufficiency (definition 8.1), the factorisation theorem 8.2, minimal sufficiency (theorem 8.6), completeness (definition 8.8), maximum and minimum (proposition 8.9). | lectures 10, 11, 13 and 26 |
| 4 | 10: Rao–Blackwell, theorem 10.1; Lehmann–Scheffe, theorem 10.4; the completeness facts, proposition 10.5; Basu, theorem 10.9, and corollary 10.10. 11: Fisher information (definition 11.2), the identities, theorem 11.3, additivity and reparametrisation (propositions 11.4 and 11.10). | lectures 13, 16, 17, 25, 26 and 29 |
| 5 | 13: the Cramer–Rao inequality, theorems 13.2 and 13.3; efficiency (definition 13.4) and attainment, proposition 13.5; the uniform counterexample, example 13.6. 14: the information matrix (definition 14.2), theorem 14.5, the normal matrix of example 14.6. | lectures 17 and 23 |
| 6 | 16: the regularity conditions (definition 16.1), theorems 16.2 and 16.3. 17: asymptotic normality, theorem 17.1; standard errors, corollary 17.3; the delta method, theorem A.18. | lectures 23, 25 and 28 |
| 7 | 19: the EM algorithm (definition 19.1), the ascent property, theorem 19.2, fixed points, proposition 19.3, the normal mixture of example 19.4. 20: tests, errors, power, size and level (definitions 20.2 to 20.5), the -value (definition 20.8), the -tests of examples 20.6 and 20.7. | lectures 22, 23, 25 and 29 |
| 8 | 22: the Neyman–Pearson lemma, theorem 22.2; one-sided optimality, proposition 22.4; no two-sided UMP test, proposition 22.5. 23: the generalised likelihood ratio (definition 23.1), Wilks’ theorem 23.3, the exact case of example 23.2. | lectures 25 and 29 |
| 9 | 25: confidence intervals and pivots (definitions 25.1 and 25.2), duality, theorem 25.6, the Wald interval, proposition 25.7. 26: posterior, theorem 26.2; conjugacy, proposition 26.4 and example 26.5; the normal–normal model, proposition 26.6; the posterior mean, theorem 26.7; Jeffreys priors, definition 26.8 and proposition 26.9. | lectures 28 and 29 |
| 10 | 28: credible and HPD intervals (definitions 28.1 and 28.2), Bayesian tests (definition 28.6, example 28.5), the comparison of section 28.3. 29: the -test and the variance test (propositions 29.1 and 29.3), table 29.1. | the examination |
| 11 | 31: the exponential model worked from start to finish, a synthesis of the whole module. |
The Results You Need to Know
For each result in this section, you must know the statement and conditions, and you must be able to reproduce the proof. A proof is reconstructed from the idea, rather than being memorised, so for each result we have written down the idea in one line. If this line is not enough for you to work out the argument, then you need to spend some time revising the proof.
We begin with estimation.
- •
-
•
Theorem 5.6, the MLE for is : a relabelling of the parameter does not change which distribution has the largest likelihood.
-
•
Theorem 7.7, the moments of the natural statistic are derivatives of : differentiate the normalising integral under the integral sign, once for the mean and twice for the variance.
-
•
Theorem 8.2, in the discrete case: write the joint probability as , and observe that sufficiency is the statement that the second factor is free of ; for the converse, sum the factorisation over the data vectors with the same value of .
- •
-
•
Theorem 10.4: two unbiased functions of a complete statistic have a difference with mean zero, and thus are equal, so that the unbiased function of is unique; Rao–Blackwell shows that no other unbiased estimator beats it.
-
•
Theorem 10.9: the conditional probability minus the unconditional one has mean zero for every , and completeness forces it to be zero.
- •
-
•
Theorems 13.2 and 13.3, the Cramer–Rao inequality: the covariance of an unbiased estimator for with the score equals , by differentiating under the integral sign, and the Cauchy–Schwarz inequality, lemma A.8, does the rest. Proposition 13.5 is the equality case of Cauchy–Schwarz: the estimator must be an affine function of the score.
- •
-
•
Theorem 16.3: on the event that exceeds both and , whose probability tends to one by the previous theorem, the log-likelihood has an interior maximum and thus, by Rolle’s theorem, a root of the likelihood equation within of ; uniqueness of the root makes it the MLE.
-
•
Theorem 17.1: expand the score about , solve for , and apply the central limit theorem to the score, the law of large numbers to the second derivative, consistency and (R4) to the remainder and Slutsky’s theorem, theorem A.16, to combine the three. Corollary 17.3 follows by replacing by in the variance, again by Slutsky’s theorem.
- •
We now turn our attention to testing, intervals and the Bayesian approach.
-
•
Theorem 22.2: The integrand is non-negative everywhere, and the theorem follows by integrating this non-negative function to get . From proposition 22.4 we know that the region of the test is the same for all alternatives. From proposition 22.5 we know that there is no two-sided uniformly most powerful test, because such a test would have to agree with both one-sided tests, which reject on opposite tails.
-
•
Theorem 25.6: The parameter is in the interval if and only if the test for does not reject the value, and thus the probability of the interval covering is one minus the size of the test.
- •
-
•
Theorem 26.2: This is an application of Bayes’ theorem, proposition A.20, where the parameter is the event and the data is the evidence. Proposition 26.4: can be multiplied by a likelihood of the form to obtain an exponent of the same shape. Theorem 26.7: The same add-and-subtract argument as in theorem 2.7, but now under the posterior distribution.
- •
- •
In addition to the theorems, the examination requires the recall of definitions and standard values, requiring only memory. The definitions are given in table 31.1 and should be given in the form in the notes, including the quantifier “for all ” wherever required. The standard values are the MLEs for the families in tables 1.1 and 1.2, as summarised in lecture 5, the Fisher information for one observation for the same families, as summarised in lecture 11, the normal information matrix (14.1) and its inverse, the one-sided and two-sided -tests and their power functions from lecture 20, and the Beta–Bernoulli and normal–normal updates from lecture 26.
What Is (Not) Proved
The following list summarises the results the module assumes to be known without proof. (No one should try to improve a non-existent proof!) For each result, the examinable material is the statement, including any conditions, and how to use the result in a calculation.
-
•
Wilks’ theorem 23.3. Know the setting (23.3), the conclusion under , and how to use the result: compute from the two maximised log-likelihoods, count as the number of constraints, and read off from the table. The argument in lecture 23 for the case is an explanation, not a proof. Also note that in example 23.2 the statement is exact for every .
-
•
The continuous case of the factorisation theorem 8.2. The proof of the discrete case is examinable, but for continuous data the theorem is used in both directions, with the indicator of a parameter-dependent support included in or as appropriate. The minimal-sufficiency criterion, theorem 8.6, is in a similar position: the proof was only sketched, but use of the result is examinable.
- •
-
•
The vector Cramer–Rao bound, theorem 14.5. Know the statement, that the difference of the covariance matrix and the inverse information is positive semi-definite, and the consequence for single coordinates, , together with the example of the normal distribution. The proof was sketched in lectures.
- •
-
•
The large-sample normal approximation (28.4) to the posterior distribution. Know the statement, that under the regularity conditions and for a prior with continuous positive density the posterior is approximately , and its consequence, that the equal-tailed credible interval is approximately the Wald interval whatever the prior. The argument given in lecture 28 is an explanation, not a proof.
The material in appendix A is background knowledge which is used without proof, but the examination expects you to name the results when they are used. All the other material from the content lectures, including the MLE asymptotics from lectures 16 and 17, and the Lehmann–Scheffe theorem, was proved in full and is examinable as such.
Writing an Answer
The model answers in appendix C show the level of detail expected. Every statistical argument in this module has the same five steps, and an answer which makes each step visible is a complete answer.
-
1.
State the model and the parameter. Write down the density or the probability weights, including the support and the parameter space, before doing anything with it. Many likelihood mistakes, in particular for models where the support depends on the parameter, are caused by leaving out the indicator of the support.
-
2.
Name the result used: “by the factorisation theorem”, “by Wilks’ theorem”. A calculation which gives the correct result without explaining how it was obtained is not a complete answer.
-
3.
Check the conditions. If the result requires a sufficient statistic to be complete, as for the Lehmann–Scheffe theorem 10.4, you should mention where the completeness comes from. If the result requires regularity conditions, as for the Cramer–Rao inequality, theorem 13.2, you should say that the support does not depend on the parameter. For a Wald interval or for Wilks’ theorem 23.3, you should say that the statement is an approximation for large samples. If a question asks why a result does not apply, this step is the whole answer.
-
4.
Compute the result. Keep the computation on the same page where you write down the argument, to allow the reader to follow your argument.
-
5.
Give the conclusion in words, in the language of the original question: “the test rejects at the level”. Use the correct kind of statement: a confidence interval is a range of values which cover the parameter with a given probability, as repeated samples are taken. A credible interval is a range of values which have a given posterior probability, given the data. A -value is not the probability that is true.
For a question like “show that” something is true, the answer is given in the question itself, and the argument is what counts. If one part of a question cannot be answered, later parts can still be answered: the questions are written so that later parts only use the statement of the previous parts, not the proofs of the previous parts.
Three habits help in every question. Distinguish the information of one observation from that of the sample, , and say which one a formula uses. Verify that a root of the likelihood equation is a maximum. Count degrees of freedom by counting constraints, and remember that a two-sided chi-squared test needs two quantiles, because the distribution is not symmetric. The next section lists the mistakes made most often.
Common Mistakes
The following mistakes are committed in scripts year after year, and they are all easy to avoid once the mistakes are known.
-
•
Treating the parameter as random in a frequentist statement. A confidence interval is a random interval which covers the fixed value with probability before the data are seen; the sentence “ lies in with probability ” is a statement about a credible interval, definition 28.1, and is false for a confidence interval, definition 25.1.
-
•
Using the Cramer–Rao inequality, or the asymptotic normality of the MLE, for a model whose support depends on the parameter. The uniform distribution on is a standard counterexample (examples 11.11 and 13.6); for such a model the estimator based on the maximum or the minimum has to be analysed directly.
-
•
Claiming that an estimator is a UMVUE because it is unbiased and a function of a sufficient statistic. The Lehmann–Scheffe theorem requires the statistic to be complete, and this completeness needs to be shown, typically using proposition 10.5.
-
•
Finding a root of the likelihood equation and calling it the MLE without checking that it is a maximum. The second derivative, or the concavity of the log-likelihood, is part of the answer, and for a model like the uniform distribution the maximum is not a root of the likelihood equation (section 5.3).
-
•
Interpreting a non-rejection as a proof of , or interpreting a -value as the probability that is true. A test controls the probability of type I errors, and the -value is a probability, computed under , about the data, definition 20.8.
-
•
Mixing up the one-observation and the sample information. The Cramer–Rao bound and the standard error use , and the limit in theorem 17.1 uses together with the factor ; both forms are equivalent, but the needs to be in one place or the other.
-
•
Leaving out the indicator of the support when factorising a density. For the uniform distribution the factor is the whole point of the factorisation, and without it the maximum would not be sufficient (exercise 8.2).
-
•
Stopping at a number. An estimate of , a statistic of or a posterior probability of is not an answer until it has been compared to the relevant threshold and turned into a sentence about the question.
Using the Practice Paper
The exercises in appendix C are the best rehearsal for the examination you will get, and it is advisable to try them yourself before looking at the solutions.
A practice paper is most useful when it is sat before the answers are seen. If you have not yet done so, sit the paper under examination conditions, with the answers written out in full. A paper which has been read through, with a mental note that each part looks manageable, tells you very little. A paper which has been sat tells you where the time went, which parts you could start but not finish, and which results you thought you knew but could not write down.
Two habits from doing the paper carry over to the examination. If you find yourself spending more than a quarter of an hour on one part of the paper, write down what you have done so far and then move on to the next question. It is better to have a partial answer for all parts of a question than to have a complete answer for only some parts.
Thank you for taking part in this module, and good luck in the examination.