Lecture 1 Statistical Models and the Estimation Problem

The subjects of probability and statistics are in some sense the opposite of each other. In a probability course we are given a distribution ℙ\mathbb{P} and we ask questions about the data generated by this distribution: how often should we expect ten tosses of a fair coin to result in seven heads? In contrast, in a statistics course we are given data and we ask questions about the distribution: what can we learn about the coin by observing seven heads?

probability:ℙ⟶X1,…,Xn,statistics:X1,…,Xn⟶ℙ.\text{probability:}\quad\mathbb{P}\;\longrightarrow\;X_{1},\dots,X_{n},\qquad% \text{statistics:}\quad X_{1},\dots,X_{n}\;\longrightarrow\;\mathbb{P}.

The topic of this module is the second arrow in the above discussion. We will consider two questions: in estimation we consider how (and how well) we can determine an unknown quantity from data. In hypothesis testing we consider whether the data are compatible with a given claim. We will answer these questions first for the classical approach, and then for the Bayesian approach. In this first lecture we will lay the foundations for these questions by introducing the model and the estimators computed from the data.

1.1 Statistical Models

The data are given as x1,…,xnx_{1},\dots,x_{n}, a list of nn numbers. Once we have computed an estimate from these numbers, we want to know how far it is from the truth, and this only makes sense if there is a true value to compare the estimate to. The data themselves do not contain a ‘true value’. Instead, this ground truth is provided by a model, a mathematical description of how the data were generated.

Our description of how the data were generated is called a model. We use a model to describe the data as the outcome of a random experiment. The data we have observed, x1,…,xnx_{1},\dots,x_{n}, are seen as the observed values of random variables X1,…,XnX_{1},\dots,X_{n}, where the distribution of the random variables depends on the unknown true value. Here we use the convention that upper case letters like XiX_{i} stand for random variables, and that lower case letters like xix_{i} stand for the observed values for our dataset.

In the simplest and most common situation the observations are a random sample: the XiX_{i} are independent and identically distributed (i.i.d.), but we do not know the distribution. Instead, we assume that the distribution belongs to a known family of distributions, labelled by a parameter.

Definition 1.1.

A parametric model for data X1,…,XnX_{1},\dots,X_{n} consists of a family

(ℙθ|θ∈Θ)\bigl{(}\mathbb{P}_{\theta}\!\mathrel{\big{|}}\!\theta\in\Theta\bigr{)}

of distributions, and the assumption that the X1,…,XnX_{1},\dots,X_{n} are i.i.d. with distribution ℙθ\mathbb{P}_{\theta} for an unknown value of the parameter θ\theta. The set Θ\Theta of possible parameter values is called the parameter space.

Our assumption here is that the actually observed data, x1,…,xnx_{1},\dots,x_{n}, are a sample from ℙθ\mathbb{P}_{\theta} for one (unknown to us) θ\theta. This value is called the true parameter value and is denoted by θ0\theta_{0}. This is the value which our estimates are compared to: an estimate is good, if it is close to θ0\theta_{0}, and bad if it is far away from θ0\theta_{0}.

Since the distribution is determined by the parameter, learning about the distribution also means learning about the parameter. Thus, the diagram from the beginning of this lecture can be updated to

probability:
θ0⟶ℙθ0⟶X1,…,Xn,\displaystyle\quad\theta_{0}\;\longrightarrow\;\mathbb{P}_{\theta_{0}}\;% \longrightarrow\;X_{1},\dots,X_{n},
statistics:
X1,…,Xn⟶θ0.\displaystyle\quad X_{1},\dots,X_{n}\;\longrightarrow\;\theta_{0}.

The true parameter value can never be observed directly, and all information about the true parameter value comes from X1,…,XnX_{1},\dots,X_{n}.

The parameter θ\theta does not necessarily have to be a single number. For example, if we assume that the data are normally distributed with unknown mean μ\mu and unknown variance σ2\sigma^{2}, then ℙθ=N⁢(μ,σ2)\mathbb{P}_{\theta}=N(\mu,\sigma^{2}) and the parameter is the pair θ=(μ,σ2)\theta=(\mu,\sigma^{2}). In this case, Θ=ℝ×(0,∞)\Theta=\mathbb{R}\times(0,\infty) is a subset of the plane. While we allow for more general parameters from the start, most of the examples in this text will involve a scalar parameter θ\theta.

The distribution ℙθ\mathbb{P}_{\theta} describes, for every set AA of possible values, the probability ℙθ⁢(A)\mathbb{P}_{\theta}(A) that an observation falls into AA. In practice, a distribution is almost always given by a function f⁢(x;θ)f(x;\theta) of the possible values xx, and this can be done in two different ways. For discrete distributions, where the observations can take finitely or countably many values, e.g. for {0,1}\{0,1\} or {0,1,2,…}\{0,1,2,\dots\}, we have f⁢(x;θ)=ℙθ⁢({x})f(x;\theta)=\mathbb{P}_{\theta}(\{x\}), giving the probability of observing the value xx. These numbers are called the probability weights of the distribution (sometimes also called the probability mass function). These weights are numbers in the interval [0,1][0,1], summing to one, and satisfying

ℙθ⁢(A)=∑x∈Af⁢(x;θ).\mathbb{P}_{\theta}(A)=\sum_{x\in A}f(x;\theta).

For continuous distributions, f⁢(x;θ)f(x;\theta) is a density: a non-negative function, which is integrable and has integral one, such that

ℙθ⁢(A)=∫Af⁢(x;θ)⁢dx.\mathbb{P}_{\theta}(A)=\int_{A}f(x;\theta)\,\mathrm{d}x.

For continuous distributions, each value xx has probability ℙθ⁢({x})=0\mathbb{P}_{\theta}\bigl{(}\{x\}\bigr{)}=0, and a density can take values larger than one. The function values f⁢(x;θ)f(x;\theta) of a density are not probabilities, but integrals of the density over sets give probabilities.

In both cases, the family of functions (f⁢(x;θ)|θ∈Θ)\bigl{(}f(x;\theta)\!\mathrel{\big{|}}\!\theta\in\Theta\bigr{)} determines the model, and we will often write the model in this form. Much of the theory in this module is common to both cases. Where a statement is true for both cases, we sometimes only state it for densities, and then the discrete case can be deduced by interpreting f⁢(x;θ)f(x;\theta) as the probability weights and all integrals over xx as sums.

We use the notation ℙθ\mathbb{P}_{\theta} also for the probabilities of events involving the data, and we write 𝔼θ\mathbb{E}_{\theta} for expectations, to indicate that the result depends on the value of the parameter. For example, ℙθ⁢(X1≤t)=ℙθ⁢((−∞,t])\mathbb{P}_{\theta}(X_{1}\leq t)=\mathbb{P}_{\theta}\bigl{(}(-\infty,t]\bigr{)} is the probability that a single observation is at most tt, when the data have distribution ℙθ\mathbb{P}_{\theta}.

We will consider a small number of families of models throughout the module. The discrete families are listed in table 1.1 and the continuous ones in table 1.2. The corresponding means and variances, which we will need frequently in calculations, are shown in tables A.1 and A.2 in appendix A.

Family Weights f⁢(x;θ)f(x;\theta) Support Θ\Theta
Bernoulli θx⁢(1−θ)1−x\theta^{x}(1-\theta)^{1-x} {0,1}\{0,1\} (0,1)(0,1)
Binomial (mx)⁢px⁢(1−p)m−x\dbinom{m}{x}p^{x}(1-p)^{m-x} {0,1,…,m}\{0,1,\dots,m\} {1,2,…}×(0,1)\{1,2,\dots\}\times(0,1)
Poisson θxx!⁢e−θ\dfrac{\theta^{x}}{x!}\,e^{-\theta} {0,1,2,…}\{0,1,2,\dots\} (0,∞)(0,\infty)
Table 1.1: Discrete families of models considered in the module. For the binomial distribution the parameter is the pair θ=(m,p)\theta=(m,p).
Family Density f⁢(x;θ)f(x;\theta) Support Θ\Theta
Exponential θ⁢e−θ⁢x\theta\,e^{-\theta x} [0,∞)[0,\infty) (0,∞)(0,\infty)
Normal 12⁢π⁢σ2⁢exp⁡(−(x−μ)22⁢σ2)\dfrac{1}{\sqrt{2\pi\sigma^{2}}}\,\exp\!\Bigl{(}-\dfrac{(x-\mu)^{2}}{2\sigma^{% 2}}\Bigr{)} ℝ\mathbb{R} ℝ×(0,∞)\mathbb{R}\times(0,\infty)
Uniform 1/θ1/\theta [0,θ][0,\theta] (0,∞)(0,\infty)
Table 1.2: Continuous families of models considered in the module. For the normal distribution the parameter is the pair θ=(μ,σ2)\theta=(\mu,\sigma^{2}).

For families where θ\theta is a vector, usually one component is assumed known and fixed, and only the remaining component is estimated. For example, for the binomial distribution the number mm of trials is often assumed to be known, and for the normal distribution the variance σ2\sigma^{2} is often assumed to be known.

Example 1.2.

Consider the model X1,…,Xn∼Uniform⁢(0,θ)X_{1},\dots,X_{n}\sim\text{Uniform}(0,\theta) with θ>0\theta>0, i.e. each observation is equally likely to take a value anywhere in the interval [0,θ][0,\theta] and f⁢(x;θ)=1/θf(x;\theta)=1/\theta for 0≤x≤θ0\leq x\leq\theta, and 0 otherwise. Since no observation can exceed θ\theta, if θ\theta increases, larger data values are possible. The range of the data can be used to draw inference about the parameter. This is in contrast to the normal distribution or the exponential distribution, where all values are possible for all θ\theta. We will see later that the moving support of the uniform distribution leads to different behaviour than for the other models, and this example will return as a test case.

1.2 Statistics and Estimators

Once a model is fixed, we want to extract information about the unknown θ\theta from the data. Since we do not know the value of θ\theta, we need to consider quantities which can be computed from the data.

Definition 1.3.

A statistic is a function T=T⁢(X1,…,Xn)T=T(X_{1},\dots,X_{n}) of the data, which does not depend on the parameter θ\theta.

Under this restriction the sample mean X¯=(X1+⋯+Xn)/n\bar{X}=(X_{1}+\dots+X_{n})/n is a statistic, but X1−θX_{1}-\theta is not, since we need to know the parameter value to compute it.

Definition 1.4.

An estimator for θ\theta is a statistic θ^=θ^⁢(X1,…,Xn)\hat{\theta}=\hat{\theta}(X_{1},\dots,X_{n}), which can be used to approximate θ\theta. The number θ^⁢(x1,…,xn)\hat{\theta}(x_{1},\dots,x_{n}), computed from a given dataset x1,…,xnx_{1},\dots,x_{n}, is called an estimate.

Throughout the notes we will use hats to denote estimators: θ^\hat{\theta} is an estimator for θ\theta, σ^2\hat{\sigma}^{2} is an estimator for σ2\sigma^{2}, and so on. If two estimators for the same quantity are compared, the second one is denoted by a tilde: θ~\tilde{\theta} is a competitor of θ^\hat{\theta}.

Most estimators are defined by a formula which works for all sample sizes, e.g. X¯=(X1+⋯+Xn)/n\bar{X}=(X_{1}+\dots+X_{n})/n. Technically, this formula defines a family of estimators (θ^n)n∈ℕ(\hat{\theta}_{n})_{n\in\mathbb{N}}, one for each sample size nn. When we need to consider the dependence on nn, and in particular when we study the behaviour of an estimator as the amount of data increases, we write θ^n\hat{\theta}_{n} instead of θ^\hat{\theta}.

An important note is that the definition does not imply anything about quality. Any statistic is an estimator, for example the constant function θ^=0\hat{\theta}=0 is a (bad) estimator for a Bernoulli success probability. We will discuss how to find and recognise good estimators in the following lectures.

The difference between an estimator and an estimate is a subtle but important one. An estimate is a single number, computed from the data, for example θ^=0.7\hat{\theta}=0.7. An estimator is a rule for computing such a number, typically a function of the random variables X1,…,XnX_{1},\dots,X_{n}, such that different outcomes of the experiment give rise to different values of the estimate.

We can only make statements about how good an estimator is, because θ^⁢(X1,…,Xn)\hat{\theta}(X_{1},\dots,X_{n}) is a random variable. The sampling distribution of an estimator is the distribution of the resulting estimates. We can ask questions like the following: If the data are generated by the true value θ0\theta_{0}, how are the estimates distributed around θ0\theta_{0}? Is the distribution centred around the truth? If so, how large is the spread? A good estimator will have a sampling distribution which is concentrated around the true parameter value, for all possible values of the true parameter. In the next lecture we will learn how to quantify these ideas using the concepts of bias and mean squared error.

1.3 The Road Ahead

Now that we have established the model and the estimator, we can start to build the module. The topics we will cover include point estimation (both likelihood and method of moments), measurement and comparison of estimators, the best possible estimators (sufficiency, information in a sample), and the limits of estimation. We will then consider statistical tests and the related question of compatibility of data with a given hypothesis about θ\theta. Finally, we will consider both estimation and testing from a Bayesian perspective, where the parameter is assumed to be random.

Summary.
  • •
    ​

    Statistics is the inverse of probability: we use data to learn about the model, rather than using the model to learn about the data.

  • •
    ​

    A parametric model assumes that the data are i.i.d. from a family (ℙθ|θ∈Θ)\bigl{(}\mathbb{P}_{\theta}\!\mathrel{\big{|}}\!\theta\in\Theta\bigr{)} of distributions, for an unknown true parameter value θ0\theta_{0}. The model provides the true value which an estimate is compared to.

  • •
    ​

    Discrete distributions are given by probability weights, which are probabilities, while continuous distributions are given by a density, which is not a probability. Probabilities are calculated as sums of weights or as integrals of the density.

  • •
    ​

    A statistic is a function of the data. An estimator is a statistic which is used to approximate θ\theta, and an estimate is the value of the estimator for a given dataset.

  • •
    ​

    An estimator is a random variable, and we study its sampling distribution. A good estimator will have a sampling distribution concentrated around the true value θ0\theta_{0}, for all possible values of θ0\theta_{0}. In the next lecture we will learn how to quantify this idea.

Exercise 1.1.

A coin with unknown probability θ\theta of heads is tossed nn times. The results of the individual tosses are given by Xi=1X_{i}=1, if the ii-th toss is heads, and Xi=0X_{i}=0 otherwise.

  1. 1.
    ​

    Write down a parametric model for X1,…,XnX_{1},\dots,X_{n}, including the probability weights and the parameter space.

  2. 2.
    ​

    Determine an estimator θ^\hat{\theta} for θ\theta.

  3. 3.
    ​

    In one sentence each, explain why θ^⁢(X1,…,Xn)\hat{\theta}(X_{1},\dots,X_{n}) is a random variable while θ^⁢(x1,…,xn)\hat{\theta}(x_{1},\dots,x_{n}) is a number.

Exercise 1.2.

Let X1,…,Xn∼N⁢(μ,σ2)X_{1},\dots,X_{n}\sim N(\mu,\sigma^{2}) be i.i.d., where both the mean μ\mu and the variance σ2\sigma^{2} are unknown, so that the parameter is θ=(μ,σ2)\theta=(\mu,\sigma^{2}). Which of the following quantities are statistics?

(a)max1≤i≤n⁡Xi−min1≤i≤n⁡Xi,(b)1n⁢∑i=1n(Xi−μ)2,(c)1n−1⁢∑i=1n(Xi−X¯)2.\text{(a)}\ \ \max_{1\leq i\leq n}X_{i}-\min_{1\leq i\leq n}X_{i},\qquad\text{% (b)}\ \ \frac{1}{n}\sum_{i=1}^{n}(X_{i}-\mu)^{2},\qquad\text{(c)}\ \ \frac{1}{% n-1}\sum_{i=1}^{n}(X_{i}-\bar{X})^{2}.
Exercise 1.3.

Let X1,…,Xn∼Uniform⁢(0,θ)X_{1},\dots,X_{n}\sim\text{Uniform}(0,\theta) be i.i.d., as in example 1.2.

  1. 1.
    ​

    Write down the joint density f⁢(x1,…,xn;θ)f(x_{1},\dots,x_{n};\theta), and state the values of the data for which it is non-zero.

  2. 2.
    ​

    Show that the estimator θ^=maxi⁡Xi\hat{\theta}=\max_{i}X_{i} always satisfies θ^≤θ\hat{\theta}\leq\theta.

  3. 3.
    ​

    Show that in fact θ^<θ\hat{\theta}<\theta with probability one, and deduce that 𝔼θ⁢(θ^)<θ\mathbb{E}_{\theta}(\hat{\theta})<\theta, so that on average the estimator underestimates θ\theta.