Introduction
These are the lecture notes for the Statistical Theory (MATH5703M) module. The question we will consider in this module, in many different forms, is the following: What can we learn about the parameters of a model if we are given data from the model?
The first half of the module is concerned with estimation. We will start by considering statistical models and properties of an estimator, using the bias and mean squared error as measures of quality. We will then consider two different approaches to constructing estimators: maximum likelihood estimators and the method of moments. The unifying framework for this part of the module is the exponential family of distributions. Sufficiency, completeness and the Rao–Blackwell and Lehmann–Scheffe theorems will be used to study how to improve an estimator and to identify cases where an estimator cannot be improved. Finally, we will consider Fisher information and the Cramer–Rao bound to study limitations on the precision of unbiased estimators. We will conclude the study of estimation by considering large-sample theory for the maximum likelihood estimator and the EM algorithm.
The second half of the module is concerned with testing and interval estimation. We will develop the Neyman–Pearson theory of hypothesis tests and the likelihood-ratio test, and will then consider confidence intervals. A section on Bayesian inference will cover the choice of prior, credible intervals and Bayesian testing. We will conclude the study of testing by considering standard tests used in practice.
Alongside the theory, we will use simulation in R to illustrate the ideas of the module. Appendix B collects the R commands we use.
How to Read the Proofs
Many of the benefits of this module can be gained by studying the proofs, but a different approach is needed to read and understand a proof than is needed to understand ordinary text. The technique we will use here, called self-explanation, is the only technique for reading proofs which is supported by strong experimental evidence. In three experiments with undergraduate mathematics students, Hodds et al. (2014) showed that a short self-study training in self-explanation, consisting of around twenty minutes of slides and one worked example and one practice proof, significantly improved scores in proof-comprehension tests. The improvement in the laboratory experiment was close to one standard deviation. A fifteen-minute booklet version of the training, completed during a scheduled lecture, also showed a smaller but significant improvement, and the effect of this was still visible three weeks later. We will practise this technique together in the first week of the module, and my hope is that you will use the technique throughout the module.
The idea of self-explanation is that you should pause after every line of the proof, and try to explain the line to yourself, before you read the next line. A good self-explanation answers the following three questions.
-
•
What does this line state?
-
•
Why is this statement true? Be specific about how you know: is it from an earlier line of the proof, or from a definition, or from an assumption of the theorem, or from a previous result?
-
•
How does this line contribute to the proof? How does it move the argument towards the statement?
There are two bad habits which look similar to self-explanation, but are not: paraphrasing, i.e. restating a line in different words without understanding the reason, and ‘monitoring’, i.e. saying ‘yes, this makes sense’ and then carrying on. Both of these techniques help your understanding less than self-explanation does. It is not enough to just have a feeling that you understand a proof; you also need to be able to reconstruct the argument line by line. If you cannot name the reason for a statement, you have a gap in your understanding and you need to find it: write down exactly what is missing, look it up or work it out, and only then continue.
In practice this means reading with a pencil. Before you read the proof, read the statement carefully, and make sure you understand what is assumed and what is claimed. Then, as you read the proof, write down the justification for each step in the margin. If you cannot find a justification for a step, try to fill the gap yourself before you read the next line. This approach will make reading a proof slower, but it will also make sure that you understand the proof. The same technique can be used to understand the worked solutions in appendix C: try to solve the exercise yourself first, and only then look at the solution line by line to check your work.
What About AI?
Large language models are very good at the kind of material we cover in this module. Using one of these models, it is possible to get a solution to any of the exercises in these notes, to get answers for the examination questions, and even to get R code for most of the tasks required for the practical report.
The problem with this is that the aim of the module is for you to learn the material in these notes, not to solve the exercises (solutions for all of the exercises are in the appendix). While AI can explain every line of a proof to you, this is not the same as self-explanation. Similarly, the practical report requires you to show your judgement and to explain your work, but a report written by the model does not give you practice in either of these skills.
There is evidence of this problem in scientific studies: Strömberg et al. (2026) studied the effect of such tools on secondary-school students over a period of 30 months: homework scores increased by 18%, homework time decreased by 30% and after six months scores in the monthly, closed-book examinations had decreased by 20%. The losses were concentrated amongst the approximately 80% of users who completed their homework faster than any non-user and scored high marks for the homework, and the authors of the study interpreted these results as indicating that these users had “outsourced” their homework to the tool. Users who spent as much time on their homework as non-users did performed about as well in the examinations, so the loss in this group was not due to the tools themselves, but due to the way the tools were used.
A randomised trial in mathematics (Bastani et al., 2025) showed the same effect. Around secondary-school students either used a language model to learn the material or used only the textbook and notes. The students were then tested in a closed-book exam. The students who used the unrestricted language model scored 17% worse than the others, mainly because they used the model to find answers to the questions. When a version of the model was used which only gave hints instead of answers, the loss in performance disappeared. The students were not aware of the difference: students who used the unrestricted model did not believe that they had learned less than the other students.
I want to also mention that the exam for our module is closed book, pen and paper based. In the exam, you will not have access to AI support, and will need to rely on what you have personally learned.
There are ways to use these models to help understanding, but these require a slightly different approach: you ask the model for something other than the answer to the question at hand.
-
•
Ask for more practice material. The model can easily generate more exercises for you. But remember that the only way to learn from these exercises is to solve them yourself!
-
•
Ask for a second explanation. If a definition, a result or a step in a proof is not clear to you, ask the model to explain it in different words; sometimes a second explanation helps you to understand it. Afterwards, close the window and write it down from memory: a definition or result in full, a step in a proof by giving the reason why it holds in your own words. Reading the model’s explanation is only monitoring, while writing it down from memory shows that you have understood or at least memorised it.
-
•
Ask which result a step relies on. If you are stuck on a step of a proof which you cannot yet justify, the model can sometimes name the result the step uses, for example Jensen’s inequality or the tower property of conditional expectation. Ask for the name only, look up the result in the index and work out the step yourself.
-
•
Ask why, not what. A question like “Why do we divide by in the sample variance?” is a question which you can first try to answer yourself, and then check your answer with the model’s answer. In contrast, a question like “solve this exercise” leaves nothing for you to do.
-
•
Ask about R. What does an error message mean? What does a given argument of a command do? These are questions which can be answered by the help pages (see appendix B). It is not expected that you will know these details by heart from this module, so asking the model is fine.
One way to test whether you have understood an argument is to try to re-create the argument on paper, from memory, a week later. If you can do this, you have learned something! If you cannot do this, you only have a solution, not an understanding.
Reading
These notes are self-contained, but the following books can be used as additional reading. The links lead to the University of Leeds Library catalogue.
-
•
G. Casella and R. L. Berger, Statistical Inference (2nd edition; available online). This is our main text. It covers most of the material of the module, from point estimation to sufficiency and the information inequality, and from hypothesis testing to confidence intervals.
-
•
G. A. Young and R. L. Smith, Essentials of Statistical Inference (print only). A concise and rigorous graduate text at the level of this module. The book is particularly good on likelihood, decision theory and the large-sample behaviour of the maximum likelihood estimator.
-
•
G. H. Givens and J. A. Hoeting, Computational Statistics (2nd edition; available online). This book is useful for learning about the EM algorithm and about numerical optimisation of the likelihood. It also introduces the computational ideas which we use in the R sessions.
-
•
P. M. Lee, Bayesian Statistics: An Introduction (4th edition; available online). A self-contained introduction to the Bayesian material we cover, including the choice of prior, Jeffreys priors and Bayesian testing.
-
•
W. J. Braun and D. J. Murdoch, A First Course in Statistical Programming with R (3rd edition; available online). This book is designed to support the R sessions and the practical report. It teaches programming and simulation in R from the ground up.
The following list has additional resources which either are more specialised and go deeper, or offer a gentler alternative to the more standard texts listed above.
-
•
D. R. Cox and D. V. Hinkley, Theoretical Statistics (available online). The classic deeper treatment of the theory behind the module.
-
•
E. L. Lehmann and G. Casella, Theory of Point Estimation (2nd edition; print only). The advanced reference for the estimation half: sufficiency, completeness and efficiency in full detail.
-
•
A. C. Davison, Statistical Models (print only). Broad and likelihood-centred, with the EM algorithm and many worked applications.
-
•
A. Gelman et al., Bayesian Data Analysis (3rd edition; available online). The standard modern reference for Bayesian statistics, well beyond what we need but the natural next step.
-
•
M. H. DeGroot and M. J. Schervish, Probability and Statistics (4th edition; available online). A gentle and thorough account, useful for revising the prerequisite material from MATH2701.
-
•
J. A. Rice, Mathematical Statistics and Data Analysis (3rd edition; print only). An accessible companion pitched a little below the core, good for a second explanation.