Lecture 1 Statistical Models and the Estimation Problem
The subjects of probability and statistics are in some sense the opposite of each other. In a probability course we are given a distribution and we ask questions about the data generated by this distribution: how often should we expect ten tosses of a fair coin to result in seven heads? In contrast, in a statistics course we are given data and we ask questions about the distribution: what can we learn about the coin by observing seven heads?
The topic of this module is the second arrow in the above discussion. We will consider two questions: in estimation we consider how (and how well) we can determine an unknown quantity from data. In hypothesis testing we consider whether the data are compatible with a given claim. We will answer these questions first for the classical approach, and then for the Bayesian approach. In this first lecture we will lay the foundations for these questions by introducing the model and the estimators computed from the data.
1.1 Statistical Models
The data are given as , a list of numbers. Once we have computed an estimate from these numbers, we want to know how far it is from the truth, and this only makes sense if there is a true value to compare the estimate to. The data themselves do not contain a ‘true value’. Instead, this ground truth is provided by a model, a mathematical description of how the data were generated.
Our description of how the data were generated is called a model. We use a model to describe the data as the outcome of a random experiment. The data we have observed, , are seen as the observed values of random variables , where the distribution of the random variables depends on the unknown true value. Here we use the convention that upper case letters like stand for random variables, and that lower case letters like stand for the observed values for our dataset.
In the simplest and most common situation the observations are a random sample: the are independent and identically distributed (i.i.d.), but we do not know the distribution. Instead, we assume that the distribution belongs to a known family of distributions, labelled by a parameter.
A parametric model for data consists of a family
of distributions, and the assumption that the are i.i.d. with distribution for an unknown value of the parameter . The set of possible parameter values is called the parameter space.
Our assumption here is that the actually observed data, , are a sample from for one (unknown to us) . This value is called the true parameter value and is denoted by . This is the value which our estimates are compared to: an estimate is good, if it is close to , and bad if it is far away from .
Since the distribution is determined by the parameter, learning about the distribution also means learning about the parameter. Thus, the diagram from the beginning of this lecture can be updated to
The true parameter value can never be observed directly, and all information about the true parameter value comes from .
The parameter does not necessarily have to be a single number. For example, if we assume that the data are normally distributed with unknown mean and unknown variance , then and the parameter is the pair . In this case, is a subset of the plane. While we allow for more general parameters from the start, most of the examples in this text will involve a scalar parameter .
The distribution describes, for every set of possible values, the probability that an observation falls into . In practice, a distribution is almost always given by a function of the possible values , and this can be done in two different ways. For discrete distributions, where the observations can take finitely or countably many values, e.g. for or , we have , giving the probability of observing the value . These numbers are called the probability weights of the distribution (sometimes also called the probability mass function). These weights are numbers in the interval , summing to one, and satisfying
For continuous distributions, is a density: a non-negative function, which is integrable and has integral one, such that
For continuous distributions, each value has probability , and a density can take values larger than one. The function values of a density are not probabilities, but integrals of the density over sets give probabilities.
In both cases, the family of functions determines the model, and we will often write the model in this form. Much of the theory in this module is common to both cases. Where a statement is true for both cases, we sometimes only state it for densities, and then the discrete case can be deduced by interpreting as the probability weights and all integrals over as sums.
We use the notation also for the probabilities of events involving the data, and we write for expectations, to indicate that the result depends on the value of the parameter. For example, is the probability that a single observation is at most , when the data have distribution .
We will consider a small number of families of models throughout the module. The discrete families are listed in table 1.1 and the continuous ones in table 1.2. The corresponding means and variances, which we will need frequently in calculations, are shown in tables A.1 and A.2 in appendix A.
| Family | Weights | Support | |
| Bernoulli | |||
| Binomial | |||
| Poisson |
| Family | Density | Support | |
| Exponential | |||
| Normal | |||
| Uniform |
For families where is a vector, usually one component is assumed known and fixed, and only the remaining component is estimated. For example, for the binomial distribution the number of trials is often assumed to be known, and for the normal distribution the variance is often assumed to be known.
Consider the model with , i.e. each observation is equally likely to take a value anywhere in the interval and for , and 0 otherwise. Since no observation can exceed , if increases, larger data values are possible. The range of the data can be used to draw inference about the parameter. This is in contrast to the normal distribution or the exponential distribution, where all values are possible for all . We will see later that the moving support of the uniform distribution leads to different behaviour than for the other models, and this example will return as a test case.
1.2 Statistics and Estimators
Once a model is fixed, we want to extract information about the unknown from the data. Since we do not know the value of , we need to consider quantities which can be computed from the data.
A statistic is a function of the data, which does not depend on the parameter .
Under this restriction the sample mean is a statistic, but is not, since we need to know the parameter value to compute it.
An estimator for is a statistic , which can be used to approximate . The number , computed from a given dataset , is called an estimate.
Throughout the notes we will use hats to denote estimators: is an estimator for , is an estimator for , and so on. If two estimators for the same quantity are compared, the second one is denoted by a tilde: is a competitor of .
Most estimators are defined by a formula which works for all sample sizes, e.g. . Technically, this formula defines a family of estimators , one for each sample size . When we need to consider the dependence on , and in particular when we study the behaviour of an estimator as the amount of data increases, we write instead of .
An important note is that the definition does not imply anything about quality. Any statistic is an estimator, for example the constant function is a (bad) estimator for a Bernoulli success probability. We will discuss how to find and recognise good estimators in the following lectures.
The difference between an estimator and an estimate is a subtle but important one. An estimate is a single number, computed from the data, for example . An estimator is a rule for computing such a number, typically a function of the random variables , such that different outcomes of the experiment give rise to different values of the estimate.
We can only make statements about how good an estimator is, because is a random variable. The sampling distribution of an estimator is the distribution of the resulting estimates. We can ask questions like the following: If the data are generated by the true value , how are the estimates distributed around ? Is the distribution centred around the truth? If so, how large is the spread? A good estimator will have a sampling distribution which is concentrated around the true parameter value, for all possible values of the true parameter. In the next lecture we will learn how to quantify these ideas using the concepts of bias and mean squared error.
1.3 The Road Ahead
Now that we have established the model and the estimator, we can start to build the module. The topics we will cover include point estimation (both likelihood and method of moments), measurement and comparison of estimators, the best possible estimators (sufficiency, information in a sample), and the limits of estimation. We will then consider statistical tests and the related question of compatibility of data with a given hypothesis about . Finally, we will consider both estimation and testing from a Bayesian perspective, where the parameter is assumed to be random.
-
•
Statistics is the inverse of probability: we use data to learn about the model, rather than using the model to learn about the data.
-
•
A parametric model assumes that the data are i.i.d. from a family of distributions, for an unknown true parameter value . The model provides the true value which an estimate is compared to.
-
•
Discrete distributions are given by probability weights, which are probabilities, while continuous distributions are given by a density, which is not a probability. Probabilities are calculated as sums of weights or as integrals of the density.
-
•
A statistic is a function of the data. An estimator is a statistic which is used to approximate , and an estimate is the value of the estimator for a given dataset.
-
•
An estimator is a random variable, and we study its sampling distribution. A good estimator will have a sampling distribution concentrated around the true value , for all possible values of . In the next lecture we will learn how to quantify this idea.
A coin with unknown probability of heads is tossed times. The results of the individual tosses are given by , if the -th toss is heads, and otherwise.
-
1.
Write down a parametric model for , including the probability weights and the parameter space.
-
2.
Determine an estimator for .
-
3.
In one sentence each, explain why is a random variable while is a number.
Let be i.i.d., where both the mean and the variance are unknown, so that the parameter is . Which of the following quantities are statistics?
Let be i.i.d., as in example 1.2.
-
1.
Write down the joint density , and state the values of the data for which it is non-zero.
-
2.
Show that the estimator always satisfies .
-
3.
Show that in fact with probability one, and deduce that , so that on average the estimator underestimates .