Appendix A Background from Probability

This appendix summarises some important results from probability which we use in the module. The prerequisite for understanding the contents of this appendix is the knowledge of the MATH2701 Statistical Methods module; all the results summarised here should be known from that module (or from an equivalent background). The purpose of this appendix is to fix notation and to provide a point of reference for the main text. None of the material in this appendix is examinable on its own. The results included here are mostly statements about fully specified distributions (known models). The opposite direction, where we try to infer an unknown parameter from data, is the one we are concerned with in the main text. We include proofs of the basic tools, because these are short and instructive. We only state the deeper limit theorems, because proofs of these results would take up valuable space and time in this text.

A.1 Expectation, Variance and Independence

We write 𝔼⁢(X)\mathbb{E}(X) for the expectation of a random variable XX and Var(X)\mathop{\mathrm{Var}}\nolimits(X) for its variance. We assume that the basic rules of algebra for expectations and variances are known. For reference, we state the following rules for expectations and variances.

Proposition A.1.

Let XX and YY be random variables with finite variance, and let aa and bb be constants. Then

𝔼⁢(a⁢X+b⁢Y)=a⁢𝔼⁢(X)+b⁢𝔼⁢(Y)andVar(a⁢X+b)=a2⁢Var(X).\mathbb{E}(aX+bY)=a\,\mathbb{E}(X)+b\,\mathbb{E}(Y)\qquad\text{and}\qquad% \mathop{\mathrm{Var}}\nolimits(aX+b)=a^{2}\,\mathop{\mathrm{Var}}\nolimits(X).

The covariance is given by Cov(X,Y)=𝔼⁢(X⁢Y)−𝔼⁢(X)⁢𝔼⁢(Y)\mathop{\mathrm{Cov}}\nolimits(X,Y)=\mathbb{E}(XY)-\mathbb{E}(X)\,\mathbb{E}(Y) and thus

Var(X+Y)=Var(X)+Var(Y)+2⁢Cov(X,Y).\mathop{\mathrm{Var}}\nolimits(X+Y)=\mathop{\mathrm{Var}}\nolimits(X)+\mathop{% \mathrm{Var}}\nolimits(Y)+2\,\mathop{\mathrm{Cov}}\nolimits(X,Y).

If XX and YY are independent, then 𝔼⁢(X⁢Y)=𝔼⁢(X)⁢𝔼⁢(Y)\mathbb{E}(XY)=\mathbb{E}(X)\,\mathbb{E}(Y) and thus the covariance is zero. Similarly, the variance of the sum can be computed as the sum of the variances.

More generally, for independent X1,…,XnX_{1},\dots,X_{n} the variance of the sum is the sum of the individual variances. We use this fact, for example, when we average a random sample. Since these identities are standard, we do not give proofs of these results here.

A.2 Standard Distributions

We will consider a small number of distributions in the module. Tables A.1 and A.2 summarise the probability weights or density, support, mean and variance of these distributions, for the discrete and for the continuous distributions respectively. The families of distributions considered in tables 1.1 and 1.2 are all included, together with their moments. The negative binomial, Pareto and Rayleigh distributions do not appear in tables 1.1 and 1.2; these distributions are used in the R sessions as a source of simulated data. The tables use the usual parameter names, but in the main text these families of distributions will be considered as parametric models where the unknown parameter is denoted by θ\theta. For example, the rate parameter λ\lambda of the exponential distribution below is the θ\theta of table 1.2. Throughout the module the exponential and gamma distributions are parametrised by their rate, as in the table and as in R’s dexp and dgamma; some textbooks use the mean 1/λ1/\lambda, or equivalently the scale, so a formula from elsewhere may need to be translated.

Family f⁢(x)f(x) Support Mean Variance
Bernoulli(p)(p) px⁢(1−p)1−xp^{x}(1-p)^{1-x} {0,1}\{0,1\} pp p⁢(1−p)p(1-p)
Binomial(m,p)(m,p) (mx)⁢px⁢(1−p)m−x\dbinom{m}{x}p^{x}(1-p)^{m-x} {0,…,m}\{0,\dots,m\} m⁢pmp m⁢p⁢(1−p)mp(1-p)
Geometric(p)(p) (1−p)x−1⁢p(1-p)^{x-1}p {1,2,…}\{1,2,\dots\} 1/p1/p (1−p)/p2(1-p)/p^{2}
Poisson(λ)(\lambda) λxx!⁢e−λ\frac{\lambda^{x}}{x!}\,e^{-\lambda} {0,1,2,…}\{0,1,2,\dots\} λ\lambda λ\lambda
Negative binomial(r,p)(r,p) (x+r−1x)⁢pr⁢(1−p)x\dbinom{x+r-1}{x}p^{r}(1-p)^{x} {0,1,2,…}\{0,1,2,\dots\} r⁢(1−p)p\frac{r(1-p)}{p} r⁢(1−p)p2\frac{r(1-p)}{p^{2}}
Table A.1: The standard discrete distributions and their moments. Here f⁢(x)f(x) gives the probability weights. For the geometric distribution, the number of trials until the first success is considered, so the support starts at 11. In contrast, some texts and the statistical software R consider the number of failures before the first success, so the support starts at 0 (see appendix B).
Family f⁢(x)f(x) Support Mean Variance
Uniform(0,θ)(0,\theta) 1/θ1/\theta [0,θ][0,\theta] θ/2\theta/2 θ2/12\theta^{2}/12
Exponential(λ)(\lambda) λ⁢e−λ⁢x\lambda\,e^{-\lambda x} [0,∞)[0,\infty) 1/λ1/\lambda 1/λ21/\lambda^{2}
Gamma(α,β)(\alpha,\beta) βαΓ⁢(α)⁢xα−1⁢e−β⁢x\frac{\beta^{\alpha}}{\Gamma(\alpha)}\,x^{\alpha-1}e^{-\beta x} (0,∞)(0,\infty) α/β\alpha/\beta α/β2\alpha/\beta^{2}
Beta(α,β)(\alpha,\beta) xα−1⁢(1−x)β−1B⁢(α,β)\frac{x^{\alpha-1}(1-x)^{\beta-1}}{B(\alpha,\beta)} [0,1][0,1] αα+β\frac{\alpha}{\alpha+\beta} α⁢β(α+β)2⁢(α+β+1)\frac{\alpha\beta}{(\alpha+\beta)^{2}(\alpha+\beta+1)}
Normal(μ,σ2)(\mu,\sigma^{2}) 12⁢π⁢σ2⁢exp⁡(−(x−μ)22⁢σ2)\frac{1}{\sqrt{2\pi\sigma^{2}}}\exp\!\Bigl{(}-\frac{(x-\mu)^{2}}{2\sigma^{2}}% \Bigr{)} ℝ\mathbb{R} μ\mu σ2\sigma^{2}
Cauchy 1π⁢(1+x2)\frac{1}{\pi(1+x^{2})} ℝ\mathbb{R} none none
Pareto(α)(\alpha) α⁢x−(α+1)\alpha\,x^{-(\alpha+1)} [1,∞)[1,\infty) αα−1\frac{\alpha}{\alpha-1} α(α−1)2⁢(α−2)\frac{\alpha}{(\alpha-1)^{2}(\alpha-2)}
Rayleigh(θ)(\theta) 2⁢xθ⁢e−x2/θ\frac{2x}{\theta}\,e^{-x^{2}/\theta} [0,∞)[0,\infty) π⁢θ/2\sqrt{\pi\theta}/2 (1−π/4)⁢θ(1-\pi/4)\,\theta
Table A.2: The standard continuous distributions and their moments. Here f⁢(x)f(x) is the density, Γ\Gamma is the gamma function and B⁢(α,β)=Γ⁢(α)⁢Γ⁢(β)/Γ⁢(α+β)B(\alpha,\beta)=\Gamma(\alpha)\Gamma(\beta)/\Gamma(\alpha+\beta) the beta function. The exponential distribution is the special case Gamma⁢(1,λ)\text{Gamma}(1,\lambda). The mean of the Pareto distribution only exists for α>1\alpha>1 and the variance only for α>2\alpha>2.

In the module we often need the distribution of a sum of independent random variables, for example when we compute the distribution of a total or of a sample mean. For several of the standard distributions the sum of independent copies again has one of the standard distributions, with only the parameters changed.

Remark.

If X1,…,XnX_{1},\dots,X_{n} are independent, then the sum of Bernoulli(p)(p) variables is Binomial(n,p)(n,p), the sum of Poisson variables is Poisson with the sum of the means, the sum of Exponential(λ)(\lambda) variables (or, more generally, of Gamma variables with the same rate) is Gamma with the sum of the shapes, and the sum of independent normals is normal with the sum of the means and the sum of the variances.

Another family, introduced in exercise 23.2 of lecture 23, is the multinomial distribution. If each of nn independent trials has one of kk possible categories, with probabilities p1,…,pkp_{1},\dots,p_{k}, where the probabilities must sum to one, then the vector of counts (N1,…,Nk)(N_{1},\dots,N_{k}) is multinomial distributed with 𝔼⁢(Nj)=n⁢pj\mathbb{E}(N_{j})=np_{j} and Var(Nj)=n⁢pj⁢(1−pj)\mathop{\mathrm{Var}}\nolimits(N_{j})=np_{j}(1-p_{j}). The individual counts are Binomial(n,pj)(n,p_{j}) and the multinomial distribution is a generalisation of the binomial distribution to the case where there are more than two categories.

A.3 Sampling from a Normal Population

The following distributions are needed to describe the sampling distributions of statistics computed from a normal sample. They form the basis of the confidence intervals in lecture 25 and of the standard tests in lecture 29. The first of these distributions also arose in the comparison of estimators for the variance in lecture 2. We start our discussion by considering the chi-squared distribution.

Definition A.2.

Let Z1,…,ZkZ_{1},\dots,Z_{k} be independent, standard normally distributed random variables. Then the distribution of Z12+⋯+Zk2Z_{1}^{2}+\dots+Z_{k}^{2} is called the chi-squared distribution with kk degrees of freedom. This distribution is denoted by χk2\chi^{2}_{k}. The mean and variance of the chi-squared distribution are kk and 2⁢k2k, respectively.

The chi-squared distribution is a special case of the gamma distribution: the χk2\chi^{2}_{k} distribution corresponds to the Gamma⁢(k/2,1/2)\text{Gamma}(k/2,1/2) distribution from table A.2, i.e. it has density xk/2−1⁢e−x/2/(2k/2⁢Γ⁢(k/2))x^{k/2-1}e^{-x/2}/\bigl{(}2^{k/2}\Gamma(k/2)\bigr{)} for x>0x>0. In statistics, the parametrisation in terms of degrees of freedom is more commonly used. The most important property of the chi-squared distribution for applications in statistics is that the sample mean and sample variance, being the two most common statistics for normally distributed samples, have known and independent distributions.

Proposition A.3.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. N⁢(μ,σ2)N(\mu,\sigma^{2}) variables. Then the sample mean X¯\bar{X} and the sample variance

S2=1n−1⁢∑i=1n(Xi−X¯)2S^{2}=\frac{1}{n-1}\sum_{i=1}^{n}(X_{i}-\bar{X})^{2}

satisfy

X¯∼N⁢(μ,σ2n),(n−1)⁢S2σ2∼χn−12,\bar{X}\sim N\Bigl{(}\mu,\ \frac{\sigma^{2}}{n}\Bigr{)},\qquad\frac{(n-1)S^{2}% }{\sigma^{2}}\sim\chi^{2}_{n-1},

and the two statistics X¯\bar{X} and S2S^{2} are independent.

The independence of X¯\bar{X} and S2S^{2} is the least obvious of these statements; we will prove this statement in lecture 10, as an application of Basu’s theorem.

We now construct two more distributions, using the normal distribution and the chi-squared distribution as building blocks. These distributions describe what happens when trying to estimate an unknown variance.

Definition A.4.

Let Z∼N⁢(0,1)Z\sim N(0,1), V∼χk2V\sim\chi^{2}_{k} and W∼χl2W\sim\chi^{2}_{l} be independent. Then the tt distribution with kk degrees of freedom is the distribution of

T=ZV/kT=\frac{Z}{\sqrt{V/k}}

denoted by tkt_{k}, and the FF distribution with kk and ll degrees of freedom is the distribution of

F=V/kW/l,F=\frac{V/k}{W/l},

denoted by Fk,lF_{k,l}.

Using definition A.4 and proposition A.3 one can understand the origin of the tt statistic for a normal sample: by dividing the standardised mean by the square root of the scaled sample variance, the unknown σ\sigma is replaced with its estimate and a tn−1t_{n-1} variable is obtained. We will see this argument in lecture 25.

A.4 Some Inequalities

A small number of basic inequalities are used throughout the module, mostly to estimate the probability of an estimator being away from its target. The most basic of these inequalities is an upper tail bound for non-negative variables, in terms of the mean.

Lemma A.5 (Markov’s inequality).

Let XX be a non-negative random variable and a>0a>0. Then

ℙ⁢(X≥a)≤𝔼⁢(X)a.\mathbb{P}(X\geq a)\leq\frac{\mathbb{E}(X)}{a}.
Proof.

By definition, XX is non-negative and thus at least aa if {X≥a}\{X\geq a\} and at least zero otherwise. Thus we can conclude X≥a⁢ 1{X≥a}X\geq a\,\mathbf{1}_{\{X\geq a\}}. Taking expectations we get

𝔼⁢(X)≥a⁢𝔼⁢(𝟏{X≥a})=a⁢ℙ⁢(X≥a).\mathbb{E}(X)\geq a\,\mathbb{E}(\mathbf{1}_{\{X\geq a\}})=a\,\mathbb{P}(X\geq a).

Dividing by the positive number aa completes the proof. ∎

The following application of Markov’s inequality, showing that a bound on the mean can be converted into a bound on the spread, completes this section.

Lemma A.6 (Chebyshev’s inequality).

Let XX have finite variance and a>0a>0. Then

ℙ⁢(|X−𝔼⁢(X)|≥a)≤Var(X)a2.\mathbb{P}\bigl{(}|X-\mathbb{E}(X)|\geq a\bigr{)}\leq\frac{\mathop{\mathrm{Var% }}\nolimits(X)}{a^{2}}.
Proof.

The event |X−𝔼⁢(X)|≥a|X-\mathbb{E}(X)|\geq a coincides with the event (X−𝔼⁢(X))2≥a2(X-\mathbb{E}(X))^{2}\geq a^{2}. By Markov’s inequality, lemma A.5, for the non-negative variable (X−𝔼⁢(X))2(X-\mathbb{E}(X))^{2} with threshold a2a^{2}, we get

ℙ⁢(|X−𝔼⁢(X)|≥a)=ℙ⁢((X−𝔼⁢(X))2≥a2)≤𝔼⁢((X−𝔼⁢(X))2)a2=Var(X)a2.\mathbb{P}\bigl{(}|X-\mathbb{E}(X)|\geq a\bigr{)}=\mathbb{P}\bigl{(}(X-\mathbb% {E}(X))^{2}\geq a^{2}\bigr{)}\leq\frac{\mathbb{E}\bigl{(}(X-\mathbb{E}(X))^{2}% \bigr{)}}{a^{2}}=\frac{\mathop{\mathrm{Var}}\nolimits(X)}{a^{2}}.

This is the required inequality. ∎

The following inequality, which establishes a relation between the expectation of a convex function and the function of the expectation, is the basis of the information inequality for maximum likelihood and of the ascent property in the EM algorithm.

Lemma A.7 (Jensen’s inequality).

Let gg be convex and XX be a random variable with finite mean. Then

g⁢(𝔼⁢(X))≤𝔼⁢(g⁢(X)).g\bigl{(}\mathbb{E}(X)\bigr{)}\leq\mathbb{E}\bigl{(}g(X)\bigr{)}.

If gg is strictly convex, then the inequality is strict unless XX is constant with probability one.

Proof.

Since gg is convex, at the point m=𝔼⁢(X)m=\mathbb{E}(X) there is a supporting line with a constant slope cc, i.e. a function g⁢(x)≥g⁢(m)+c⁢(x−m)g(x)\geq g(m)+c\,(x-m) for all xx. Substituting XX for xx and taking expectations, the linear term has expectation zero since 𝔼⁢(X−m)=0\mathbb{E}(X-m)=0, and thus

𝔼⁢(g⁢(X))≥g⁢(m)+c⁢𝔼⁢(X−m)=g⁢(m)=g⁢(𝔼⁢(X)).\mathbb{E}\bigl{(}g(X)\bigr{)}\geq g(m)+c\,\mathbb{E}(X-m)=g(m)=g\bigl{(}% \mathbb{E}(X)\bigr{)}.

If gg is strictly convex, the supporting line at mm is unique and thus g⁢(x)>g⁢(m)+c⁢(x−m)g(x)>g(m)+c\,(x-m) for all x≠mx\neq m. In this case, if we have equality 𝔼⁢(g⁢(X))=g⁢(m)\mathbb{E}\bigl{(}g(X)\bigr{)}=g(m), the non-negative random variable g⁢(X)−g⁢(m)−c⁢(X−m)g(X)-g(m)-c\,(X-m) has expectation zero and thus is zero with probability one, i.e. X=mX=m with probability one. This proves the claimed inequality and completes the proof. ∎

Finally, the Cauchy–Schwarz inequality can be used to estimate a covariance in terms of variances. This result is, for example, used in the derivation of the Cramer–Rao lower bound in lecture 13.

Lemma A.8 (Cauchy–Schwarz inequality).

Let XX and YY be random variables with finite variances. Then we have

Cov(X,Y)2≤Var(X)⁢Var(Y).\mathop{\mathrm{Cov}}\nolimits(X,Y)^{2}\leq\mathop{\mathrm{Var}}\nolimits(X)\,% \mathop{\mathrm{Var}}\nolimits(Y).
Proof.

We can assume that both variances are positive, since otherwise one of the random variables is constant and both sides of the inequality equal zero. Let U=X−𝔼⁢(X)U=X-\mathbb{E}(X) and V=Y−𝔼⁢(Y)V=Y-\mathbb{E}(Y), so that Cov(X,Y)=𝔼⁢(U⁢V)\mathop{\mathrm{Cov}}\nolimits(X,Y)=\mathbb{E}(UV), Var(X)=𝔼⁢(U2)\mathop{\mathrm{Var}}\nolimits(X)=\mathbb{E}(U^{2}) and Var(Y)=𝔼⁢(V2)\mathop{\mathrm{Var}}\nolimits(Y)=\mathbb{E}(V^{2}), and we want to show 𝔼⁢(U⁢V)2≤𝔼⁢(U2)⁢𝔼⁢(V2)\mathbb{E}(UV)^{2}\leq\mathbb{E}(U^{2})\,\mathbb{E}(V^{2}). For any real number tt, the quantity 𝔼⁢((U−t⁢V)2)\mathbb{E}\bigl{(}(U-tV)^{2}\bigr{)} is an expectation of a square and thus is non-negative. This gives

𝔼⁢(U2)−2⁢t⁢𝔼⁢(U⁢V)+t2⁢𝔼⁢(V2)≥0\mathbb{E}(U^{2})-2t\,\mathbb{E}(UV)+t^{2}\,\mathbb{E}(V^{2})\geq 0

for all tt. A quadratic polynomial which never takes a negative value has non-positive discriminant and thus we have 4⁢𝔼⁢(U⁢V)2−4⁢𝔼⁢(U2)⁢𝔼⁢(V2)≤04\,\mathbb{E}(UV)^{2}-4\,\mathbb{E}(U^{2})\,\mathbb{E}(V^{2})\leq 0 and thus we get the required inequality. This completes the proof. ∎

A.5 Conditional Expectation

When one random variable carries information about another, we summarise it through conditional expectation. Given random variables XX and YY, the conditional expectation 𝔼⁢(X|Y)\mathbb{E}(X\mskip 1.0mu|\mskip 1.0muY) is the expectation of XX computed as if the value of YY were known; because that value is itself random, 𝔼⁢(X|Y)\mathbb{E}(X\mskip 1.0mu|\mskip 1.0muY) is a function of YY, and hence a random variable in its own right. Its own average recovers the ordinary expectation of XX.

Proposition A.9 (Tower property).

Let XX and YY be random variables, where XX has finite mean. Then

𝔼⁢(𝔼⁢(X|Y))=𝔼⁢(X).\mathbb{E}\bigl{(}\mathbb{E}(X\mskip 1.0mu|\mskip 1.0muY)\bigr{)}=\mathbb{E}(X).

Since the value of 𝔼⁢(X|Y)\mathbb{E}(X\mskip 1.0mu|\mskip 1.0muY) is computed from the value of YY, we can pull out all factors which are functions of YY from the conditional expectation.

Proposition A.10 (Taking out what is known).

Let XX and YY be random variables and let hh be a function such that h⁢(Y)⁢Xh(Y)X has finite mean. Then

𝔼⁢(h⁢(Y)⁢X|Y)=h⁢(Y)⁢𝔼⁢(X|Y).\mathbb{E}\bigl{(}h(Y)\,X\!\mathrel{\big{|}}\!Y\bigr{)}=h(Y)\,\mathbb{E}(X% \mskip 1.0mu|\mskip 1.0muY).

In particular 𝔼⁢(h⁢(Y)|Y)=h⁢(Y)\mathbb{E}\bigl{(}h(Y)\!\mathrel{\big{|}}\!Y\bigr{)}=h(Y).

We already know the tower property and this rule. The only effect of conditioning can be to reduce the variance on average. The exact argument is given by the following decomposition, which we will use in lecture 10 when we improve an estimator by conditioning.

Proposition A.11 (Law of total variance).

Let XX be a random variable with finite variance and let YY be any random variable. Then

Var(X)=𝔼⁢(Var(X|Y))+Var(𝔼⁢(X|Y)),\mathop{\mathrm{Var}}\nolimits(X)=\mathbb{E}\bigl{(}\mathop{\mathrm{Var}}% \nolimits(X\mskip 1.0mu|\mskip 1.0muY)\bigr{)}+\mathop{\mathrm{Var}}\nolimits% \bigl{(}\mathbb{E}(X\mskip 1.0mu|\mskip 1.0muY)\bigr{)},

where Var(X|Y)=𝔼⁢(X2|Y)−𝔼⁢(X|Y)2\mathop{\mathrm{Var}}\nolimits(X\mskip 1.0mu|\mskip 1.0muY)=\mathbb{E}\bigl{(}% X^{2}\!\mathrel{\big{|}}\!Y\bigr{)}-\mathbb{E}(X\mskip 1.0mu|\mskip 1.0muY)^{2} denotes the conditional variance.

Proof.

Let m=𝔼⁢(X|Y)m=\mathbb{E}(X\mskip 1.0mu|\mskip 1.0muY) for the conditional mean. Using the tower property, proposition A.9, for X2X^{2} we find 𝔼⁢(X2)=𝔼⁢(𝔼⁢(X2|Y))\mathbb{E}(X^{2})=\mathbb{E}\bigl{(}\mathbb{E}\bigl{(}X^{2}\!\mathrel{\big{|}}% \!Y\bigr{)}\bigr{)} and using the definition of the conditional variance we have 𝔼⁢(X2|Y)=Var(X|Y)+m2\mathbb{E}\bigl{(}X^{2}\!\mathrel{\big{|}}\!Y\bigr{)}=\mathop{\mathrm{Var}}% \nolimits(X\mskip 1.0mu|\mskip 1.0muY)+m^{2}. Taking expectations we get

𝔼⁢(X2)=𝔼⁢(Var(X|Y))+𝔼⁢(m2).\mathbb{E}(X^{2})=\mathbb{E}\bigl{(}\mathop{\mathrm{Var}}\nolimits(X\mskip 1.0% mu|\mskip 1.0muY)\bigr{)}+\mathbb{E}(m^{2}).

Using the tower property again, we know 𝔼⁢(m)=𝔼⁢(X)\mathbb{E}(m)=\mathbb{E}(X) and thus, after subtracting 𝔼⁢(X)2=𝔼⁢(m)2\mathbb{E}(X)^{2}=\mathbb{E}(m)^{2} on both sides, we can write 𝔼⁢(X2)−𝔼⁢(X)2\mathbb{E}(X^{2})-\mathbb{E}(X)^{2} as Var(X)\mathop{\mathrm{Var}}\nolimits(X) on the left-hand side and 𝔼⁢(m2)−𝔼⁢(m)2\mathbb{E}(m^{2})-\mathbb{E}(m)^{2} as Var(m)\mathop{\mathrm{Var}}\nolimits(m) on the right-hand side. This is the required identity and completes the proof. ∎

A.6 Convergence and Limit Theorems

The large-sample theory of the module is based on two different concepts of a sequence of random variables converging to a limit. The weaker of the two concepts requires that the random variables are unlikely to differ from the limit by more than a fixed amount.

Definition A.12.

A sequence of random variables XnX_{n} converges in probability to a limit XX, written as Xn→pXX_{n}\xrightarrow{\ \mathrm{p}\ }X, if for every ε>0\varepsilon>0

ℙ⁢(|Xn−X|>ε)⟶0(n→∞).\mathbb{P}\bigl{(}|X_{n}-X|>\varepsilon\bigr{)}\longrightarrow 0\qquad(n\to% \infty).

The weaker (but still useful) convergence of distribution functions is used to approximate a whole distribution:

Definition A.13.

A sequence of random variables XnX_{n} with distribution functions FnF_{n} converges in distribution to a limit XX with distribution function FF, written as Xn→dXX_{n}\xrightarrow{\ \mathrm{d}\ }X, if Fn⁢(x)→F⁢(x)F_{n}(x)\to F(x) for all points xx where FF is continuous.

The two great theorems of large-sample statistics describe the behaviour of a sample mean of independent, identically distributed observations. The first theorem gives the mean of the sample, the second describes the size and shape of the fluctuations around the mean.

Theorem A.14 (Weak law of large numbers).

Let X1,X2,…X_{1},X_{2},\dots be i.i.d. with finite mean μ\mu. Then the sample mean X¯n\bar{X}_{n} satisfies X¯n→pμ\bar{X}_{n}\xrightarrow{\ \mathrm{p}\ }\mu as n→∞n\to\infty.

The assumption of finite mean is important. The Cauchy distribution from table A.2 has no finite mean and the distribution of the sample mean is the same as the distribution of a single observation for all nn. Thus, in this case nothing converges and this example is used in lecture 2 to argue that the sample mean is not consistent.

Theorem A.15 (Central limit theorem).

Let X1,X2,…X_{1},X_{2},\dots be i.i.d. with mean μ\mu and finite variance σ2\sigma^{2}. Then

n⁢(X¯n−μ)→dN⁢(0,σ2)(n→∞).\sqrt{n}\,(\bar{X}_{n}-\mu)\xrightarrow{\ \mathrm{d}\ }N(0,\sigma^{2})\qquad(n% \to\infty).

In order to combine these limits we will need rules which allow us to manipulate limits. Slutsky’s theorem allows us to treat a term which converges to a constant as the constant itself, and the continuous mapping theorem allows us to apply a continuous function to a limit.

Theorem A.16 (Slutsky’s theorem).

Let Xn→dXX_{n}\xrightarrow{\ \mathrm{d}\ }X and Yn→pcY_{n}\xrightarrow{\ \mathrm{p}\ }c for some constant cc. Then Xn+Yn→dX+cX_{n}+Y_{n}\xrightarrow{\ \mathrm{d}\ }X+c and Xn⁢Yn→dc⁢XX_{n}Y_{n}\xrightarrow{\ \mathrm{d}\ }cX. If c≠0c\neq 0 we have Xn/Yn→dX/cX_{n}/Y_{n}\xrightarrow{\ \mathrm{d}\ }X/c.

Theorem A.17 (Continuous mapping theorem).

Let gg be continuous at the possible values of XX. Then Xn→dXX_{n}\xrightarrow{\ \mathrm{d}\ }X implies g⁢(Xn)→dg⁢(X)g(X_{n})\xrightarrow{\ \mathrm{d}\ }g(X) and Xn→pXX_{n}\xrightarrow{\ \mathrm{p}\ }X implies g⁢(Xn)→pg⁢(X)g(X_{n})\xrightarrow{\ \mathrm{p}\ }g(X).

An important application of this theorem is to apply the central limit theorem after applying a smooth transformation to the data, in order to determine the limiting distribution of a function of an asymptotically normal statistic. This method is known as the delta method. We will use the delta method in lecture 17 to compute standard errors for functions of the maximum likelihood estimator.

Theorem A.18 (Delta method).

Let n⁢(Tn−θ)→dN⁢(0,τ2)\sqrt{n}\,(T_{n}-\theta)\xrightarrow{\ \mathrm{d}\ }N(0,\tau^{2}) and let gg be a differentiable function at θ\theta with g′⁢(θ)≠0g^{\prime}(\theta)\neq 0. Then

n⁢(g⁢(Tn)−g⁢(θ))→dN⁢(0,g′⁢(θ)2⁢τ2).\sqrt{n}\,\bigl{(}g(T_{n})-g(\theta)\bigr{)}\xrightarrow{\ \mathrm{d}\ }N\bigl% {(}0,\ g^{\prime}(\theta)^{2}\,\tau^{2}\bigr{)}.

These five theorems are given without proof, and are from the material covered in a first course in probability. We use the theorems as a ready-made toolkit.

A.7 Transformation of a Random Variable

We sometimes know the distribution of a random variable and want to find the distribution of a smooth, one-to-one transformation of this random variable. The change-of-variable formula is a mathematical tool which allows us to find the density of the transformed random variable. This formula is, for example, used when we reparametrise a model, as discussed in the context of the Jeffreys prior in lecture 26.

Proposition A.19 (Change of variable).

Let XX be a continuous random variable with density fXf_{X} and let Y=g⁢(X)Y=g(X) for a differentiable, strictly monotone function gg with inverse hh and with g′⁢(x)≠0g^{\prime}(x)\neq 0 for all xx. Then YY has density

fY⁢(y)=fX⁢(h⁢(y))⁢|h′⁢(y)|f_{Y}(y)=f_{X}\bigl{(}h(y)\bigr{)}\,\bigl{|}h^{\prime}(y)\bigr{|}

for all yy in the range of gg.

Proof.

We only consider the case where gg is increasing, the proof for the decreasing case is similar with changes in sign. The distribution function of YY is given by

FY⁢(y)=ℙ⁢(g⁢(X)≤y)=ℙ⁢(X≤h⁢(y))=FX⁢(h⁢(y)),F_{Y}(y)=\mathbb{P}(g(X)\leq y)=\mathbb{P}\bigl{(}X\leq h(y)\bigr{)}=F_{X}% \bigl{(}h(y)\bigr{)},

where we used that gg is increasing and has inverse hh. Taking derivatives with respect to yy, using the chain rule, we find fY⁢(y)=fX⁢(h⁢(y))⁢h′⁢(y)f_{Y}(y)=f_{X}(h(y))\,h^{\prime}(y) and thus h′⁢(y)>0h^{\prime}(y)>0 equals its own absolute value as given in the formula. This completes the proof. ∎

The same argument applies to a vector of random variables: if XX has a joint density fXf_{X} on ℝk\mathbb{R}^{k} and Y=g⁢(X)Y=g(X) for a differentiable, one-to-one map gg with inverse hh, then YY has joint density fY⁢(y)=fX⁢(h⁢(y))⁢|deth′⁢(y)|f_{Y}(y)=f_{X}\bigl{(}h(y)\bigr{)}\,|\det h^{\prime}(y)|, where h′⁢(y)h^{\prime}(y) is the Jacobian matrix of hh. For a linear map with determinant one, as in the solution to exercise 10.2 in appendix C, the factor is 11.

A.8 Bayes’ Theorem

The final tool reverses the direction of conditioning, expressing the conditional density of one variable given another in terms of the reverse conditional. It is the mechanism behind every posterior distribution in the Bayesian lectures.

Proposition A.20 (Bayes’ theorem).

Let AA and BB be random variables with joint density f⁢(a,b)f(a,b) and let BB have strictly positive marginal density f⁢(b)f(b). Then the conditional density of AA given B=bB=b is

f⁢(a|b)=f⁢(b|a)⁢f⁢(a)f⁢(b),f(a\mskip 1.0mu|\mskip 1.0mub)=\frac{f(b\mskip 1.0mu|\mskip 1.0mua)\,f(a)}{f(b% )},

where f⁢(b)=∫f⁢(b|a)⁢f⁢(a)⁢daf(b)=\int f(b\mskip 1.0mu|\mskip 1.0mua)\,f(a)\,\mathrm{d}a is the marginal density of BB.

If one of the two variables is discrete, the same formula holds with its probability weights in place of the density, and the integral in f⁢(b)f(b) becomes a sum over the values of AA when AA is discrete.

In the Bayesian lectures of the module, the unknown parameter plays the role of AA with prior density f⁢(a)f(a) and the data are given by BB. In this situation, proposition A.20 states that the posterior density is proportional to the product of the likelihood and the prior density. The denominator f⁢(B)f(B) does not depend on the parameter. This viewpoint is further developed in lectures 26 and 28.