Appendix A Background from Probability
This appendix summarises some important results from probability which we use in the module. The prerequisite for understanding the contents of this appendix is the knowledge of the MATH2701 Statistical Methods module; all the results summarised here should be known from that module (or from an equivalent background). The purpose of this appendix is to fix notation and to provide a point of reference for the main text. None of the material in this appendix is examinable on its own. The results included here are mostly statements about fully specified distributions (known models). The opposite direction, where we try to infer an unknown parameter from data, is the one we are concerned with in the main text. We include proofs of the basic tools, because these are short and instructive. We only state the deeper limit theorems, because proofs of these results would take up valuable space and time in this text.
A.1 Expectation, Variance and Independence
We write for the expectation of a random variable and for its variance. We assume that the basic rules of algebra for expectations and variances are known. For reference, we state the following rules for expectations and variances.
Let and be random variables with finite variance, and let and be constants. Then
The covariance is given by and thus
If and are independent, then and thus the covariance is zero. Similarly, the variance of the sum can be computed as the sum of the variances.
More generally, for independent the variance of the sum is the sum of the individual variances. We use this fact, for example, when we average a random sample. Since these identities are standard, we do not give proofs of these results here.
A.2 Standard Distributions
We will consider a small number of distributions in the module. Tables A.1 and A.2 summarise the probability weights or density, support, mean and variance of these distributions, for the discrete and for the continuous distributions respectively. The families of distributions considered in tables 1.1 and 1.2 are all included, together with their moments. The negative binomial, Pareto and Rayleigh distributions do not appear in tables 1.1 and 1.2; these distributions are used in the R sessions as a source of simulated data. The tables use the usual parameter names, but in the main text these families of distributions will be considered as parametric models where the unknown parameter is denoted by . For example, the rate parameter of the exponential distribution below is the of table 1.2. Throughout the module the exponential and gamma distributions are parametrised by their rate, as in the table and as in R’s dexp and dgamma; some textbooks use the mean , or equivalently the scale, so a formula from elsewhere may need to be translated.
| Family | Support | Mean | Variance | |
| Bernoulli | ||||
| Binomial | ||||
| Geometric | ||||
| Poisson | ||||
| Negative binomial |
| Family | Support | Mean | Variance | |
| Uniform | ||||
| Exponential | ||||
| Gamma | ||||
| Beta | ||||
| Normal | ||||
| Cauchy | none | none | ||
| Pareto | ||||
| Rayleigh |
In the module we often need the distribution of a sum of independent random variables, for example when we compute the distribution of a total or of a sample mean. For several of the standard distributions the sum of independent copies again has one of the standard distributions, with only the parameters changed.
If are independent, then the sum of Bernoulli variables is Binomial, the sum of Poisson variables is Poisson with the sum of the means, the sum of Exponential variables (or, more generally, of Gamma variables with the same rate) is Gamma with the sum of the shapes, and the sum of independent normals is normal with the sum of the means and the sum of the variances.
Another family, introduced in exercise 23.2 of lecture 23, is the multinomial distribution. If each of independent trials has one of possible categories, with probabilities , where the probabilities must sum to one, then the vector of counts is multinomial distributed with and . The individual counts are Binomial and the multinomial distribution is a generalisation of the binomial distribution to the case where there are more than two categories.
A.3 Sampling from a Normal Population
The following distributions are needed to describe the sampling distributions of statistics computed from a normal sample. They form the basis of the confidence intervals in lecture 25 and of the standard tests in lecture 29. The first of these distributions also arose in the comparison of estimators for the variance in lecture 2. We start our discussion by considering the chi-squared distribution.
Let be independent, standard normally distributed random variables. Then the distribution of is called the chi-squared distribution with degrees of freedom. This distribution is denoted by . The mean and variance of the chi-squared distribution are and , respectively.
The chi-squared distribution is a special case of the gamma distribution: the distribution corresponds to the distribution from table A.2, i.e. it has density for . In statistics, the parametrisation in terms of degrees of freedom is more commonly used. The most important property of the chi-squared distribution for applications in statistics is that the sample mean and sample variance, being the two most common statistics for normally distributed samples, have known and independent distributions.
Let be i.i.d. variables. Then the sample mean and the sample variance
satisfy
and the two statistics and are independent.
The independence of and is the least obvious of these statements; we will prove this statement in lecture 10, as an application of Basu’s theorem.
We now construct two more distributions, using the normal distribution and the chi-squared distribution as building blocks. These distributions describe what happens when trying to estimate an unknown variance.
Let , and be independent. Then the distribution with degrees of freedom is the distribution of
denoted by , and the distribution with and degrees of freedom is the distribution of
denoted by .
Using definition A.4 and proposition A.3 one can understand the origin of the statistic for a normal sample: by dividing the standardised mean by the square root of the scaled sample variance, the unknown is replaced with its estimate and a variable is obtained. We will see this argument in lecture 25.
A.4 Some Inequalities
A small number of basic inequalities are used throughout the module, mostly to estimate the probability of an estimator being away from its target. The most basic of these inequalities is an upper tail bound for non-negative variables, in terms of the mean.
Let be a non-negative random variable and . Then
By definition, is non-negative and thus at least if and at least zero otherwise. Thus we can conclude . Taking expectations we get
Dividing by the positive number completes the proof. ∎
The following application of Markov’s inequality, showing that a bound on the mean can be converted into a bound on the spread, completes this section.
Let have finite variance and . Then
The event coincides with the event . By Markov’s inequality, lemma A.5, for the non-negative variable with threshold , we get
This is the required inequality. ∎
The following inequality, which establishes a relation between the expectation of a convex function and the function of the expectation, is the basis of the information inequality for maximum likelihood and of the ascent property in the EM algorithm.
Let be convex and be a random variable with finite mean. Then
If is strictly convex, then the inequality is strict unless is constant with probability one.
Since is convex, at the point there is a supporting line with a constant slope , i.e. a function for all . Substituting for and taking expectations, the linear term has expectation zero since , and thus
If is strictly convex, the supporting line at is unique and thus for all . In this case, if we have equality , the non-negative random variable has expectation zero and thus is zero with probability one, i.e. with probability one. This proves the claimed inequality and completes the proof. ∎
Finally, the Cauchy–Schwarz inequality can be used to estimate a covariance in terms of variances. This result is, for example, used in the derivation of the Cramer–Rao lower bound in lecture 13.
Let and be random variables with finite variances. Then we have
We can assume that both variances are positive, since otherwise one of the random variables is constant and both sides of the inequality equal zero. Let and , so that , and , and we want to show . For any real number , the quantity is an expectation of a square and thus is non-negative. This gives
for all . A quadratic polynomial which never takes a negative value has non-positive discriminant and thus we have and thus we get the required inequality. This completes the proof. ∎
A.5 Conditional Expectation
When one random variable carries information about another, we summarise it through conditional expectation. Given random variables and , the conditional expectation is the expectation of computed as if the value of were known; because that value is itself random, is a function of , and hence a random variable in its own right. Its own average recovers the ordinary expectation of .
Let and be random variables, where has finite mean. Then
Since the value of is computed from the value of , we can pull out all factors which are functions of from the conditional expectation.
Let and be random variables and let be a function such that has finite mean. Then
In particular .
We already know the tower property and this rule. The only effect of conditioning can be to reduce the variance on average. The exact argument is given by the following decomposition, which we will use in lecture 10 when we improve an estimator by conditioning.
Let be a random variable with finite variance and let be any random variable. Then
where denotes the conditional variance.
Let for the conditional mean. Using the tower property, proposition A.9, for we find and using the definition of the conditional variance we have . Taking expectations we get
Using the tower property again, we know and thus, after subtracting on both sides, we can write as on the left-hand side and as on the right-hand side. This is the required identity and completes the proof. ∎
A.6 Convergence and Limit Theorems
The large-sample theory of the module is based on two different concepts of a sequence of random variables converging to a limit. The weaker of the two concepts requires that the random variables are unlikely to differ from the limit by more than a fixed amount.
A sequence of random variables converges in probability to a limit , written as , if for every
The weaker (but still useful) convergence of distribution functions is used to approximate a whole distribution:
A sequence of random variables with distribution functions converges in distribution to a limit with distribution function , written as , if for all points where is continuous.
The two great theorems of large-sample statistics describe the behaviour of a sample mean of independent, identically distributed observations. The first theorem gives the mean of the sample, the second describes the size and shape of the fluctuations around the mean.
Let be i.i.d. with finite mean . Then the sample mean satisfies as .
The assumption of finite mean is important. The Cauchy distribution from table A.2 has no finite mean and the distribution of the sample mean is the same as the distribution of a single observation for all . Thus, in this case nothing converges and this example is used in lecture 2 to argue that the sample mean is not consistent.
Let be i.i.d. with mean and finite variance . Then
In order to combine these limits we will need rules which allow us to manipulate limits. Slutsky’s theorem allows us to treat a term which converges to a constant as the constant itself, and the continuous mapping theorem allows us to apply a continuous function to a limit.
Let and for some constant . Then and . If we have .
Let be continuous at the possible values of . Then implies and implies .
An important application of this theorem is to apply the central limit theorem after applying a smooth transformation to the data, in order to determine the limiting distribution of a function of an asymptotically normal statistic. This method is known as the delta method. We will use the delta method in lecture 17 to compute standard errors for functions of the maximum likelihood estimator.
Let and let be a differentiable function at with . Then
These five theorems are given without proof, and are from the material covered in a first course in probability. We use the theorems as a ready-made toolkit.
A.7 Transformation of a Random Variable
We sometimes know the distribution of a random variable and want to find the distribution of a smooth, one-to-one transformation of this random variable. The change-of-variable formula is a mathematical tool which allows us to find the density of the transformed random variable. This formula is, for example, used when we reparametrise a model, as discussed in the context of the Jeffreys prior in lecture 26.
Let be a continuous random variable with density and let for a differentiable, strictly monotone function with inverse and with for all . Then has density
for all in the range of .
We only consider the case where is increasing, the proof for the decreasing case is similar with changes in sign. The distribution function of is given by
where we used that is increasing and has inverse . Taking derivatives with respect to , using the chain rule, we find and thus equals its own absolute value as given in the formula. This completes the proof. ∎
The same argument applies to a vector of random variables: if has a joint density on and for a differentiable, one-to-one map with inverse , then has joint density , where is the Jacobian matrix of . For a linear map with determinant one, as in the solution to exercise 10.2 in appendix C, the factor is .
A.8 Bayes’ Theorem
The final tool reverses the direction of conditioning, expressing the conditional density of one variable given another in terms of the reverse conditional. It is the mechanism behind every posterior distribution in the Bayesian lectures.
Let and be random variables with joint density and let have strictly positive marginal density . Then the conditional density of given is
where is the marginal density of .
If one of the two variables is discrete, the same formula holds with its probability weights in place of the density, and the integral in becomes a sum over the values of when is discrete.
In the Bayesian lectures of the module, the unknown parameter plays the role of with prior density and the data are given by . In this situation, proposition A.20 states that the posterior density is proportional to the product of the likelihood and the prior density. The denominator does not depend on the parameter. This viewpoint is further developed in lectures 26 and 28.