Preparing for the Exam

This chapter summarises what you need to do to prepare for the examination: it includes a map of the module, a summary of important results, information about what you were required to prove and what you were only required to know, advice on how to write your answers, details of common mistakes to avoid, and information on how to use the practice paper. None of this is new material, but it is gathered together here from different parts of the module for revision purposes.

The Module in One Page

The module consists of one question, split into two parts. The first part, covering lectures 1 to 19, deals with estimation: the question is what an estimator is, and how one can measure quality of an estimator. In this part of the module we also learn how estimators can be constructed using the method of moments and the maximum likelihood method, how a sufficient statistic can be used to compress the data and how conditioning on the data can be used to improve an estimator. In this part of the module we also learn about the Fisher information and how it can be used to give a bound on the variance of all unbiased estimators. We learn how the maximum likelihood estimator behaves for large sample size and how to compute the maximum likelihood estimator when some of the data are missing. Throughout this part of the module we will consider the exponential families, introduced in lecture 7. These families of distributions provide the sufficient statistic, the completeness of the sufficient statistic, the Fisher information and, in lecture 26, the conjugate priors. The second part of the module, covering lectures 20 to 29, deals with testing and interval estimation. In this part of the module we learn about the language of tests, the best test for simple hypotheses, the likelihood ratio for composite hypotheses, the duality of tests and confidence intervals, the Bayesian approach to the same questions and the standard tests for normal data. Table 31.1 shows the lectures for each week of the module and the results of each lecture and the later lectures which depend on the results of these lectures.

Week Lectures and main results Used later in
1 1: model, statistic, estimator (definitions 1.1, 1.3 and 1.4). 2: bias and mean squared error, theorem 2.7; consistency, theorem 2.10. every later lecture; the bias corrections of lecture 5; the consistency proof of lecture 16
2 4: likelihood, score and method of moments (definitions 4.1, 4.4 and 4.6). 5: the MLE (definition 5.1), the likelihood equation and its check, invariance, theorem 5.6; failures, section 5.3. lectures 11, 16, 17, 19 and 23
3 7: exponential families (definition 7.1), the cumulant function, theorem 7.7, full rank (definition 7.10). 8: sufficiency (definition 8.1), the factorisation theorem 8.2, minimal sufficiency (theorem 8.6), completeness (definition 8.8), maximum and minimum (proposition 8.9). lectures 10, 11, 13 and 26
4 10: Rao–Blackwell, theorem 10.1; Lehmann–Scheffe, theorem 10.4; the completeness facts, proposition 10.5; Basu, theorem 10.9, and corollary 10.10. 11: Fisher information (definition 11.2), the identities, theorem 11.3, additivity and reparametrisation (propositions 11.4 and 11.10). lectures 13, 16, 17, 25, 26 and 29
5 13: the Cramer–Rao inequality, theorems 13.2 and 13.3; efficiency (definition 13.4) and attainment, proposition 13.5; the uniform counterexample, example 13.6. 14: the information matrix (definition 14.2), theorem 14.5, the normal matrix of example 14.6. lectures 17 and 23
6 16: the regularity conditions (definition 16.1), theorems 16.2 and 16.3. 17: asymptotic normality, theorem 17.1; standard errors, corollary 17.3; the delta method, theorem A.18. lectures 23, 25 and 28
7 19: the EM algorithm (definition 19.1), the ascent property, theorem 19.2, fixed points, proposition 19.3, the normal mixture of example 19.4. 20: tests, errors, power, size and level (definitions 20.2 to 20.5), the pp-value (definition 20.8), the zz-tests of examples 20.6 and 20.7. lectures 22, 23, 25 and 29
8 22: the Neyman–Pearson lemma, theorem 22.2; one-sided optimality, proposition 22.4; no two-sided UMP test, proposition 22.5. 23: the generalised likelihood ratio (definition 23.1), Wilks’ theorem 23.3, the exact case of example 23.2. lectures 25 and 29
9 25: confidence intervals and pivots (definitions 25.1 and 25.2), duality, theorem 25.6, the Wald interval, proposition 25.7. 26: posterior, theorem 26.2; conjugacy, proposition 26.4 and example 26.5; the normal–normal model, proposition 26.6; the posterior mean, theorem 26.7; Jeffreys priors, definition 26.8 and proposition 26.9. lectures 28 and 29
10 28: credible and HPD intervals (definitions 28.1 and 28.2), Bayesian tests (definition 28.6, example 28.5), the comparison of section 28.3. 29: the tt-test and the variance test (propositions 29.1 and 29.3), table 29.1. the examination
11 31: the exponential model worked from start to finish, a synthesis of the whole module.
Table 31.1: The lectures week by week, with the results each contributes and the lectures which build on them. The R sessions are omitted.

The Results You Need to Know

For each result in this section, you must know the statement and conditions, and you must be able to reproduce the proof. A proof is reconstructed from the idea, rather than being memorised, so for each result we have written down the idea in one line. If this line is not enough for you to work out the argument, then you need to spend some time revising the proof.

We begin with estimation.

  • •
    ​

    Theorem 2.7, the decomposition MSE=Var+bias2\mathop{\mathrm{MSE}}\nolimits=\mathop{\mathrm{Var}}\nolimits+\mathop{\mathrm{% bias}}\nolimits^{2}: add and subtract the mean inside the square, and the cross term vanishes. Theorem 2.10, a vanishing mean squared error gives consistency: Markov’s inequality, lemma A.5, applied to (θ^n−θ)2(\hat{\theta}_{n}-\theta)^{2}.

  • •
    ​

    Theorem 5.6, the MLE for g⁢(θ)g(\theta) is g⁢(θ^)g(\hat{\theta}): a relabelling of the parameter does not change which distribution has the largest likelihood.

  • •
    ​

    Theorem 7.7, the moments of the natural statistic are derivatives of KK: differentiate the normalising integral under the integral sign, once for the mean and twice for the variance.

  • •
    ​

    Theorem 8.2, in the discrete case: write the joint probability as ℙθ⁢(T=t)⁢ℙθ⁢(X=x|T=t)\mathbb{P}_{\theta}(T=t)\,\mathbb{P}_{\theta}(X=x\mskip 1.0mu|\mskip 1.0muT=t), and observe that sufficiency is the statement that the second factor is free of θ\theta; for the converse, sum the factorisation over the data vectors with the same value of TT.

  • •
    ​

    Theorem 10.1: the conditional expectation given a sufficient statistic is a statistic, the tower property keeps it unbiased, and the law of total variance, proposition A.11, shows that the variance can only decrease.

  • •
    ​

    Theorem 10.4: two unbiased functions of a complete statistic have a difference with mean zero, and thus are equal, so that the unbiased function of TT is unique; Rao–Blackwell shows that no other unbiased estimator beats it.

  • •
    ​

    Theorem 10.9: the conditional probability ℙθ⁢(A∈B|T)\mathbb{P}_{\theta}(A\in B\mskip 1.0mu|\mskip 1.0muT) minus the unconditional one has mean zero for every θ\theta, and completeness forces it to be zero.

  • •
    ​

    Theorem 11.3: differentiate ∫f⁢(x;θ)⁢dx=1\int f(x;\theta)\,\mathrm{d}x=1 once to see that the score has mean zero, lemma 11.1, and a second time, using the quotient rule, to obtain ℐ︀n=−𝔼θ⁢(ℓ′′)\mathcal{I}_{n}=-\mathbb{E}_{\theta}(\ell^{\prime\prime}).

  • •
    ​

    Theorems 13.2 and 13.3, the Cramer–Rao inequality: the covariance of an unbiased estimator for g⁢(θ)g(\theta) with the score equals g′⁢(θ)g^{\prime}(\theta), by differentiating 𝔼θ⁢(W)=g⁢(θ)\mathbb{E}_{\theta}(W)=g(\theta) under the integral sign, and the Cauchy–Schwarz inequality, lemma A.8, does the rest. Proposition 13.5 is the equality case of Cauchy–Schwarz: the estimator must be an affine function of the score.

  • •
    ​

    Theorem 16.2: by the law of large numbers the average log-likelihood ratio converges to

    𝔼θ0⁢(log⁡(f⁢(X;θ)/f⁢(X;θ0))),\mathbb{E}_{\theta_{0}}\bigl{(}\log(f(X;\theta)/f(X;\theta_{0}))\bigr{)},

    and Jensen’s inequality, lemma A.7, makes this expectation negative.

  • •
    ​

    Theorem 16.3: on the event that L⁢(θ0)L(\theta_{0}) exceeds both L⁢(θ0−ε)L(\theta_{0}-\varepsilon) and L⁢(θ0+ε)L(\theta_{0}+\varepsilon), whose probability tends to one by the previous theorem, the log-likelihood has an interior maximum and thus, by Rolle’s theorem, a root of the likelihood equation within ε\varepsilon of θ0\theta_{0}; uniqueness of the root makes it the MLE.

  • •
    ​

    Theorem 17.1: expand the score about θ0\theta_{0}, solve for n⁢(θ^n−θ0)\sqrt{n}(\hat{\theta}_{n}-\theta_{0}), and apply the central limit theorem to the score, the law of large numbers to the second derivative, consistency and (R4) to the remainder and Slutsky’s theorem, theorem A.16, to combine the three. Corollary 17.3 follows by replacing θ0\theta_{0} by θ^n\hat{\theta}_{n} in the variance, again by Slutsky’s theorem.

  • •
    ​

    Theorem 19.2: write ℓ=Q−H\ell=Q-H as in (19.1); the M step increases QQ, and Jensen’s inequality shows that H⁢(θ|θk)H(\theta\mskip 1.0mu|\mskip 1.0mu\theta_{k}) is largest at θ=θk\theta=\theta_{k}, so that HH cannot increase.

We now turn our attention to testing, intervals and the Bayesian approach.

  • •
    ​

    Theorem 22.2: The integrand (𝟏C−𝟏C′)⁢(f⁢(x;θ1)−k⁢f⁢(x;θ0))(\mathbf{1}_{C}-\mathbf{1}_{C^{\prime}})\bigl{(}f(x;\theta_{1})-k\,f(x;\theta_% {0})\bigr{)} is non-negative everywhere, and the theorem follows by integrating this non-negative function to get βC⁢(θ1)−βC′⁢(θ1)≥k⁢(βC⁢(θ0)−βC′⁢(θ0))≥0\beta_{C}(\theta_{1})-\beta_{C^{\prime}}(\theta_{1})\geq k\bigl{(}\beta_{C}(% \theta_{0})-\beta_{C^{\prime}}(\theta_{0})\bigr{)}\geq 0. From proposition 22.4 we know that the region of the test is the same for all alternatives. From proposition 22.5 we know that there is no two-sided uniformly most powerful test, because such a test would have to agree with both one-sided tests, which reject on opposite tails.

  • •
    ​

    Theorem 25.6: The parameter θ0\theta_{0} is in the interval C⁢(X)C(X) if and only if the test for θ0\theta_{0} does not reject the value, and thus the probability of the interval covering θ0\theta_{0} is one minus the size of the test.

  • •
    ​

    Proposition 25.7: From corollary 17.3 we know that the standardised error is approximately standard normally distributed, and so bracketing it between ±z1−α/2\pm z_{1-\alpha/2} gives the interval.

  • •
    ​

    Theorem 26.2: This is an application of Bayes’ theorem, proposition A.20, where the parameter is the event and the data is the evidence. Proposition 26.4: exp⁡(a⁢η−b⁢K⁢(η))\exp(a\eta-bK(\eta)) can be multiplied by a likelihood of the form exp⁡(η⁢Tn−n⁢K⁢(η))\exp(\eta T_{n}-nK(\eta)) to obtain an exponent of the same shape. Theorem 26.7: The same add-and-subtract argument as in theorem 2.7, but now under the posterior distribution.

  • •
    ​

    Proposition 26.9: The change of variables for densities, proposition A.19, provides the factor |h′⁢(ψ)||h^{\prime}(\psi)| and the reparametrisation rule for the information, proposition 11.10, provides the same factor under the square root.

  • •
    ​

    Propositions 29.1 and 29.3: The statistics are pivots from lecture 25, evaluated at the null value, and the null distribution is given by proposition A.3 and definition A.4.

In addition to the theorems, the examination requires the recall of definitions and standard values, requiring only memory. The definitions are given in table 31.1 and should be given in the form in the notes, including the quantifier “for all θ∈Θ\theta\in\Theta” wherever required. The standard values are the MLEs for the families in tables 1.1 and 1.2, as summarised in lecture 5, the Fisher information for one observation for the same families, as summarised in lecture 11, the normal information matrix (14.1) and its inverse, the one-sided and two-sided zz-tests and their power functions from lecture 20, and the Beta–Bernoulli and normal–normal updates from lecture 26.

What Is (Not) Proved

The following list summarises the results the module assumes to be known without proof. (No one should try to improve a non-existent proof!) For each result, the examinable material is the statement, including any conditions, and how to use the result in a calculation.

  • •
    ​

    Wilks’ theorem 23.3. Know the setting (23.3), the conclusion Wn→dχr2W_{n}\xrightarrow{\ \mathrm{d}\ }\chi^{2}_{r} under H0H_{0}, and how to use the result: compute WW from the two maximised log-likelihoods, count rr as the number of constraints, and read off χr,1−α2\chi^{2}_{r,1-\alpha} from the table. The argument in lecture 23 for the case k=r=1k=r=1 is an explanation, not a proof. Also note that in example 23.2 the statement is exact for every nn.

  • •
    ​

    The continuous case of the factorisation theorem 8.2. The proof of the discrete case is examinable, but for continuous data the theorem is used in both directions, with the indicator of a parameter-dependent support included in gg or hh as appropriate. The minimal-sufficiency criterion, theorem 8.6, is in a similar position: the proof was only sketched, but use of the result is examinable.

  • •
    ​

    The completeness results, proposition 10.5. Both results can be quoted as needed, and in the case of the exponential family the full-rank condition from definition 7.10 must be mentioned. Any other completeness results will be provided in the question.

  • •
    ​

    The vector Cramer–Rao bound, theorem 14.5. Know the statement, that the difference of the covariance matrix and the inverse information is positive semi-definite, and the consequence for single coordinates, Varθ(Tj)≥(ℐ︀n−1)j⁢j\mathop{\mathrm{Var}}\nolimits_{\theta}(T_{j})\geq(\mathcal{I}_{n}^{-1})_{jj}, together with the example of the normal distribution. The proof was sketched in lectures.

  • •
    ​

    The vector case of theorem 17.1, as given at the end of lecture 17: the MLE of a vector parameter is asymptotically normal with covariance matrix ℐ︀n⁢(θ0)−1\mathcal{I}_{n}(\theta_{0})^{-1}. Know the statement and what it means for the normal distribution, where the diagonal matrix shows the separation of the two coordinates.

  • •
    ​

    The large-sample normal approximation (28.4) to the posterior distribution. Know the statement, that under the regularity conditions and for a prior with continuous positive density the posterior is approximately N⁢(θ^n,1/ℐ︀n⁢(θ^n))N\bigl{(}\hat{\theta}_{n},1/\mathcal{I}_{n}(\hat{\theta}_{n})\bigr{)}, and its consequence, that the equal-tailed credible interval is approximately the Wald interval whatever the prior. The argument given in lecture 28 is an explanation, not a proof.

The material in appendix A is background knowledge which is used without proof, but the examination expects you to name the results when they are used. All the other material from the content lectures, including the MLE asymptotics from lectures 16 and 17, and the Lehmann–Scheffe theorem, was proved in full and is examinable as such.

Writing an Answer

The model answers in appendix C show the level of detail expected. Every statistical argument in this module has the same five steps, and an answer which makes each step visible is a complete answer.

  1. 1.
    ​

    State the model and the parameter. Write down the density or the probability weights, including the support and the parameter space, before doing anything with it. Many likelihood mistakes, in particular for models where the support depends on the parameter, are caused by leaving out the indicator of the support.

  2. 2.
    ​

    Name the result used: “by the factorisation theorem”, “by Wilks’ theorem”. A calculation which gives the correct result without explaining how it was obtained is not a complete answer.

  3. 3.
    ​

    Check the conditions. If the result requires a sufficient statistic to be complete, as for the Lehmann–Scheffe theorem 10.4, you should mention where the completeness comes from. If the result requires regularity conditions, as for the Cramer–Rao inequality, theorem 13.2, you should say that the support does not depend on the parameter. For a Wald interval or for Wilks’ theorem 23.3, you should say that the statement is an approximation for large samples. If a question asks why a result does not apply, this step is the whole answer.

  4. 4.
    ​

    Compute the result. Keep the computation on the same page where you write down the argument, to allow the reader to follow your argument.

  5. 5.
    ​

    Give the conclusion in words, in the language of the original question: “the test rejects H0H_{0} at the 5%5\% level”. Use the correct kind of statement: a confidence interval is a range of values which cover the parameter with a given probability, as repeated samples are taken. A credible interval is a range of values which have a given posterior probability, given the data. A pp-value is not the probability that H0H_{0} is true.

For a question like “show that” something is true, the answer is given in the question itself, and the argument is what counts. If one part of a question cannot be answered, later parts can still be answered: the questions are written so that later parts only use the statement of the previous parts, not the proofs of the previous parts.

Three habits help in every question. Distinguish the information of one observation from that of the sample, ℐ︀n=n⁢ℐ︀\mathcal{I}_{n}=n\mathcal{I}, and say which one a formula uses. Verify that a root of the likelihood equation is a maximum. Count degrees of freedom by counting constraints, and remember that a two-sided chi-squared test needs two quantiles, because the distribution is not symmetric. The next section lists the mistakes made most often.

Common Mistakes

The following mistakes are committed in scripts year after year, and they are all easy to avoid once the mistakes are known.

  • •
    ​

    Treating the parameter as random in a frequentist statement. A confidence interval is a random interval which covers the fixed value θ\theta with probability 0.950.95 before the data are seen; the sentence “θ\theta lies in [l,u][l,u] with probability 0.950.95” is a statement about a credible interval, definition 28.1, and is false for a confidence interval, definition 25.1.

  • •
    ​

    Using the Cramer–Rao inequality, or the asymptotic normality of the MLE, for a model whose support depends on the parameter. The uniform distribution on (0,θ)(0,\theta) is a standard counterexample (examples 11.11 and 13.6); for such a model the estimator based on the maximum or the minimum has to be analysed directly.

  • •
    ​

    Claiming that an estimator is a UMVUE because it is unbiased and a function of a sufficient statistic. The Lehmann–Scheffe theorem requires the statistic to be complete, and this completeness needs to be shown, typically using proposition 10.5.

  • •
    ​

    Finding a root of the likelihood equation and calling it the MLE without checking that it is a maximum. The second derivative, or the concavity of the log-likelihood, is part of the answer, and for a model like the uniform distribution the maximum is not a root of the likelihood equation (section 5.3).

  • •
    ​

    Interpreting a non-rejection as a proof of H0H_{0}, or interpreting a pp-value as the probability that H0H_{0} is true. A test controls the probability of type I errors, and the pp-value is a probability, computed under H0H_{0}, about the data, definition 20.8.

  • •
    ​

    Mixing up the one-observation and the sample information. The Cramer–Rao bound and the standard error use ℐ︀n⁢(θ)=n⁢ℐ︀⁢(θ)\mathcal{I}_{n}(\theta)=n\mathcal{I}(\theta), and the limit in theorem 17.1 uses ℐ︀⁢(θ)\mathcal{I}(\theta) together with the factor n\sqrt{n}; both forms are equivalent, but the nn needs to be in one place or the other.

  • •
    ​

    Leaving out the indicator of the support when factorising a density. For the uniform distribution the factor 𝟏{x(n)≤θ}\mathbf{1}_{\{x_{(n)}\leq\theta\}} is the whole point of the factorisation, and without it the maximum would not be sufficient (exercise 8.2).

  • •
    ​

    Stopping at a number. An estimate of 0.1110.111, a statistic of 3.613.61 or a posterior probability of 0.570.57 is not an answer until it has been compared to the relevant threshold and turned into a sentence about the question.

Using the Practice Paper

The exercises in appendix C are the best rehearsal for the examination you will get, and it is advisable to try them yourself before looking at the solutions.

A practice paper is most useful when it is sat before the answers are seen. If you have not yet done so, sit the paper under examination conditions, with the answers written out in full. A paper which has been read through, with a mental note that each part looks manageable, tells you very little. A paper which has been sat tells you where the time went, which parts you could start but not finish, and which results you thought you knew but could not write down.

Two habits from doing the paper carry over to the examination. If you find yourself spending more than a quarter of an hour on one part of the paper, write down what you have done so far and then move on to the next question. It is better to have a partial answer for all parts of a question than to have a complete answer for only some parts.

Thank you for taking part in this module, and good luck in the examination.