Lecture 20 Hypothesis Testing
So far we have asked questions about the value of . In this lecture we will consider the second question from lecture 1: if there is a claim about , is it supported by the data? We will spend some time to carefully set up the language of hypothesis testing, which will be used in lectures 22, 23, 25 and 29. We will conclude the lecture by considering two examples, both about the mean of a normal sample.
20.1 Hypotheses
Throughout, follow a parametric model in the sense of definition 1.1. Questions like the one about fairness of a coin or about the effect of a treatment ask whether the true parameter value belongs to a subset of the parameter space .
A hypothesis is a statement of the form for a subset . In a testing problem, the null hypothesis and the alternative hypothesis are given, where and are disjoint subsets of . A hypothesis is called simple, if the corresponding subset consists of a single point, and composite otherwise.
As an abbreviation, from this lecture on we write for the value chosen to test a null hypothesis. In lectures 16 and 17, the same symbol was used to denote the unknown, true value of the parameter. The context should make it clear which meaning is intended.
In most problems we have , i.e. we can assume that exactly one of the two hypotheses is true. For example, for the mean of a normal sample, the hypothesis is simple, whereas the one-sided alternative , the two-sided alternative and the one-sided null hypothesis are all composite.
The two hypotheses have different roles: the null hypothesis is the default position, and the alternative is the claim which requires evidence before we can act on it. Section 20.3 will make this distinction precise.
20.2 Tests, Critical Regions and Errors
A test decides between the two hypotheses under consideration, and can be described by the set of data values where it rejects .
A test of against is a rule which for every possible data vector decides whether to “reject ” or “do not reject ”. The set
is called the critical region (or rejection region) of the test. Alternatively, a test can be described by its test function for and otherwise, where indicates rejection.
In practice, the critical region is almost always of the form or for a statistic given by definition 1.3, called the test statistic, and a constant called the critical value. The two forms of the critical region differ if has positive probability (exercise 20.2 shows an example of this case, for discrete distributions).
A test applied to random data will sometimes wrongly decide. This can happen in two different ways.
A type I error occurs, if the test wrongly rejects when is true. A type II error occurs, if the test wrongly does not reject when is true.
Table 20.1 summarises the possible outcomes of statistical tests, in relation to the truth of the hypothesis being tested.
| do not reject | reject | |
| true | correct decision | type I error |
| true | type II error | correct decision (power) |
To understand the interpretation of this table, we can use the following example: in a criminal trial, the null hypothesis is the presumption of innocence of the defendant, and the alternative is guilt. If an innocent person is convicted, this is a type I error. If a guilty person is acquitted, this is a type II error. The standard of proof “beyond reasonable doubt” is such that many acquittals of the guilty are acceptable in order to avoid too many wrongful convictions. The reason for this is that the two types of error have different consequences: the consequences of type I errors are more serious. Furthermore, since an acquittal does not prove innocence, we write “do not reject ” instead of “accept ” in the table.
20.3 The Power Function, Size and Level
The probabilities of the two errors depend on the true parameter value, and both can be found using a single function of .
The power function of a test with critical region is given by
i.e. the probability of rejecting when the true parameter value is . For the value is called the power of the test at the alternative .
For the value is the probability of a type I error, and for the value is the probability of a type II error. The two probabilities cannot both be reduced by making the critical region larger or smaller. If the size of the critical region is increased, the value increases for all , i.e. type I errors become more likely and type II errors become less likely. If the size of the critical region is decreased, the opposite happens. For this reason, we have to choose one of the two types of error to control and it transpires that we can decide which type of error to control by considering the asymmetry between the two types of hypotheses: type I errors are errors where the default position is abandoned without good reason, and thus we should restrict the probability of type I errors first and only then should we consider the power of the test.
The size of a test with power function is given by
where the supremum is the maximum probability of type I error among all possible null hypotheses. For , the test is a level test, if the size of the test is less than or equal to .
For a simple null hypothesis , the size of the test equals . The number is called the significance level and the values and are only conventions. Among all level tests, we prefer the test with the largest power at the alternatives. In lecture 22 we will see how to find the test with optimal power, when both the null and the alternative hypothesis are simple.
20.4 Testing the Mean of a Normal Sample
We now construct a test from first principles for i.i.d. , where the variance is known. Here we write for the standard normal distribution function and for the -quantile of the standard normal distribution, i.e. we have . The most commonly used values of this function are and .
We want to test against at significance level . If is significantly larger than , we can conclude that the alternative is more likely to be true. Thus, we can choose as the test statistic and reject if . From proposition A.3 we know that under the standardised mean is standard normally distributed. Thus, the probability of type I errors is
Setting this probability equal to , we find the critical value as and thus the level test rejects if
This is the one-sided -test. Since is standard normally distributed under , we can use the same standardisation to determine the power of the test. The power function is
Since the function is increasing, the graph of is an S-shaped curve which starts at and converges to , passing through the point . Close to the power is only slightly larger than , but as increases, the curve gets steeper. This fact is used in exercise 20.1 to determine the required sample size.
For , , and the test rejects if . At the alternative the power is , at it is only : a shift of half a standard deviation is detected in barely a third of all samples.
Since the power function (20.2) is increasing, the same test also has size for the composite null hypothesis , attained at the boundary point; this is why definition 20.5 takes a supremum.
Assume that we want to test for both directions, i.e. we test against and we reject for . Under the test statistic is standard normally distributed and, since the normal distribution is symmetric, we have
Setting this equal to , we find and thus the level two-sided -test rejects if
Let . Then, similar to the result in example 20.6, we find the power function
This function is symmetric about and has its minimum value at this point. Thus, the graph of this power function is U-shaped instead of S-shaped. For , , and the test rejects if and at we have and the power .
At the one-sided test has power against , since it rejects all values with probability from the upper tail. Thus, the one-sided test is superior to the two-sided test for all alternatives above , but is useless for alternatives below. Before we consider the data, we need to decide which departures from are of interest. In lecture 22 we will see that there is no single test which is best for both directions, and in lecture 24 we will consider how size and power behave in simulation.
20.5 The -Value
A test at a fixed level only states whether is rejected or not. The -value, introduced in this section, allows us to quantify the strength of the evidence in a more detailed way.
Consider a family of tests with critical regions for all levels , such that whenever . Then the -value of the data is given by
the smallest level at which the data lead to rejection of .
The tests from section 20.4 are nested in this way, since decreases as increases. For a test which rejects for large values of a statistic and which has a simple null hypothesis, the smallest level at which the data can lead to rejection is
This is the probability under of the test statistic taking a value at least as extreme as the observed value. For the one-sided -test this probability is , for the observed value . For the two-sided test we get . For the situation of example 20.6, the observed mean leads to and thus to for the one-sided test and for the two-sided test. We can see that we can reject at level if and only if , and thus the -value is a compact summary of a test. Under , the -value of a continuous test statistic is uniformly distributed on (see exercise 20.4). This result is used in lecture 30.
Both statements assume that the test statistic is continuous. For discrete test statistics, e.g. the number of successes in a Bernoulli sample, the tail probability (20.5) can only take the countably many values and the -value is discontinuous. Most levels are not attained by any critical region. We can still perform the test for and the resulting test is conservative. The size of the test is the largest attainable value which does not exceed the level . The test will reject a true less often than the nominal level suggests and the test will lose some power. The -value is no longer uniformly distributed either, but still satisfies . This is what allows the level to still be valid.
Since the -value is a probability, it is tempting to interpret the results in wrong ways; the following statements about a -value of are all wrong.
- •
-
•
“The probability that the observed data could have been produced by chance is .” The -value corresponds to extreme values of , not to the probability of the data.
-
•
“If we had observed a -value of , we would have known that is true.” Data which is compatible with may be compatible with many different alternatives: the test from example 20.6 has power only at .
-
•
“The effect is important, because is small.” If is large enough, any difference between and will result in a tiny -value. Reports should include both the estimate and the standard error.
None of these mistakes invalidates the -value as a tool, but they do show that the -value answers a very specific question, and that the same value should not be used to answer different questions.
-
•
A testing problem contrasts a null hypothesis against an alternative . The hypothesis is simple, if the corresponding subset of is a single point, and composite otherwise.
-
•
A test is given by its critical region , typically of the form for a test statistic and critical value .
-
•
Type I errors are cases where is true but is wrongly rejected by the test, type II errors are cases where is false but is not (correctly) rejected. The power function gives the probability of type I errors for and the power of the test for .
-
•
The size of a test is . A level test has size at most . The level is chosen first, and then the power of the test can be considered. For example, for the normal mean with known variance, the one-sided and two-sided -tests can be derived.
-
•
The -value is the smallest level at which the data lead to rejection of the hypothesis. Note that this is not the probability that is true, and a large -value does not establish .
Consider the one-sided -test from example 20.6, at level , and let be the alternative and be the target power.
-
1.
Show that the test has power at least at if and only if
-
2.
For , , and target power (so that ), determine the smallest sample size which achieves the target power.
-
3.
How does the required sample size change, when the difference is halved? What happens when is doubled?
Let be i.i.d. Bernoulli with success probability and consider testing against using the test which rejects if for an integer .
-
1.
Write down the size of the test as a function of , using the binomial distribution of under .
-
2.
Determine the smallest such that the test has level . Compute the size of the test for this case. Justify your answer by showing that no choice of can lead to a size of exactly .
-
3.
Determine the power of the test from the previous part when , and . Comment on the capabilities and limitations of the test with ten observations.
-
4.
Nine successes are observed. Compute the -value and state the decision at level .
Let be i.i.d. as in example 1.2 and let , where the density was found in example 8.10. We test against .
-
1.
Show that for .
-
2.
Consider the test which rejects if , for a constant . Determine the size of this test and the value of which makes the size equal to .
-
3.
Determine the power function of the test from the previous part, and evaluate this function at , , and .
-
4.
Consider the second test, which rejects if . Show that this test has size . Determine the power function of this test and compare the power functions of the two tests. Which of the two tests would you choose and why?
A sample of size from with known has mean . We want to test .
-
1.
Determine the -value of the one-sided -test with against and state the decisions for and .
-
2.
Determine the -value of the two-sided test against and state the same two decisions.
-
3.
Your colleague summarises the first part of the question as “there is a chance that the mean is ”. What is wrong with this sentence? Write one correct sentence to summarise the first part of the question.
-
4.
Show that under the -value of the one-sided -test, with , satisfies for all .