Lecture 29 Standard Tests Revisited
The most commonly used statistical tests, the -test, the -test, the chi-squared test for a variance and the -test, are often taught as recipes. In this lecture we will consider these tests again, from the perspective of the module: each of these tests is based on a pivot in the sense of lecture 25, and the test is performed by comparing a statistic to a quantile from a table. We will work through the one-sample -test in detail, including how to read the tables, and will state the variance, paired, two-sample and -tests in a similar way. We will conclude by giving a table which shows which test is appropriate for which question.
29.1 Tests from Pivots
Every test in this lecture is based on the following principle: the test statistic is known to follow a certain distribution under the null hypothesis, and the test rejects if the statistic exceeds a quantile of this distribution, taken from a table. Using the language of lecture 25, the test statistic is a pivot for the hypothesised value and the test is the one which, when inverted by theorem 25.6, gives the confidence interval based on this pivot.
We use the notation , and for the quantiles from lecture 25, so that for . A table shows these quantiles for a few values of . For a one-sided test at level , we need the -quantile, and for a two-sided test at the same level we need the -quantile. A -value in the sense of definition 20.8 cannot be found exactly in a table, but can be bracketed between two tabulated levels; we will see an example of this approach in example 29.2, below.
The most basic example is the -test from lecture 20: if are i.i.d. with known , then is standard normally distributed under and in examples 20.6 and 20.7 we can reject for or for the one-sided and two-sided alternative, respectively. The simplicity of this test is countered by the fact that the variance of the data is known. The rest of the lecture is concerned with how to replace when is unknown.
29.2 The One-Sample -Test
Let be i.i.d. , where both parameters are unknown, and let and be the sample mean and sample variance, respectively. If we replace in the definition of by the estimate , we get the test statistic
In example 25.4 we have seen that for this is the ratio of the standard normal variable to the square root of , where by proposition A.3, and that the two values are independent of each other by corollary 10.10. Consequently, by definition A.4. This leads to the following tests.
Let and let be i.i.d. , where the parameters and are unknown, and let be given by (29.1). Then the three tests of , which reject if
for the alternatives , and , respectively, have size , independent of the value of .
Under we have for every value of , as shown above. By the definition of the quantile we have . This is the size of the first test. The density of the distribution is symmetric about zero and thus we have for the second test and for the third test we have . Since these probabilities do not depend on , the proof is complete. ∎
These are the one-sided and two-sided one-sample -tests. Since for every , the -test requires a larger standardised deviation than the -test before it can reject the hypothesis, and this is the price we pay for estimating . As increases, this price decreases as the distribution gets closer to the standard normal distribution. As in lecture 20, the one-sided tests have level for the composite null hypotheses and , since the probability of rejection increases as moves into the alternative.
A quantity has nominal value , and ten measurements give
Assume that the measurements are i.i.d. . We want to test against at level . The sum of the observations is , so that ; the sum of the squared deviations from is , so that and . The observed value of the test statistic is
The row corresponding to degrees of freedom in the -table shows that we have , , and . For a two-sided test with level we need and since we can reject : the data are not compatible with a mean of . The same row also brackets the -value. The observed value falls between and and thus the two-sided -value, for , is between and ; the two-sided test also rejects at level but not at level . If instead the question had been whether the measurements were high, the one-sided test for against at level would have compared to and would have rejected , with a -value between and .
The -test is the test obtained by applying the likelihood-ratio principle from lecture 23. In exercise 23.1 the generalised likelihood-ratio statistic for , where is unknown for both hypotheses, is found to be , which is a strictly increasing function of . Thus, rejecting large values of corresponds to the two-sided -test. Following the general method of lecture 23, we would use the critical value from the approximation of Wilks’ theorem, theorem 23.3; however, in this case the null distribution of is known exactly and the quantile is the exact critical value for every . As gets large, the two approaches will be equal, since and and then .
29.3 The Chi-Squared Test for a Variance
The second question we can ask a normal sample, about the spread of the sample, is whether the variance equals a given value . In proposition A.3 we have seen that the pivot does not depend on the mean, and the test will follow the structure of the test from the previous section.
Let and let be i.i.d. , where and are unknown, and let
Each of the three tests for , which reject if
for the alternatives , and , respectively, has size , whatever the value of .
Under we have by proposition A.3, for every value of . The first two rejection probabilities are and by the definition of the quantiles. For the third, the events and are disjoint and each have probability , so the rejection probability is . This completes the proof. ∎
Unlike the distribution, the chi-squared distribution is not symmetric and so we need different values from the table for the two tails of a two-sided test. This is the inverse of the confidence interval from example 25.5. Similarly, the one-sided test against has level for the composite null hypothesis , since the probability of the test to reject the null hypothesis increases with . This is exercise 29.2.
The numerical carrying-out of the test is also part of that exercise, and it follows the pattern of example 29.2 with the chi-squared table in place of the -table. Note that tests about a variance have low power for small samples, because is a much more variable estimator than .
29.4 Paired and Two-Sample Problems
The tests so far have considered a single sample. Here we will consider three more standard tests, which compare two sets of measurements each, and are again pivot tests.
The first of these three tests is the paired -test. Assume that we have observed data in pairs , , e.g. before and after measurements for the same subject. Since the two measurements for one subject are not independent, we can reduce the data for each pair to the difference , model the differences as i.i.d. and then apply proposition 29.1 with : the test statistic is , constructed from the sample mean and variance of the differences, and is distributed under . The model is only for the differences, as the example in exercise 29.3 shows.
The second test is the two-sample -test. Here we assume that are i.i.d. and that are i.i.d. , where the two samples are independent of each other, and where we assume that the unknown variances of the two populations are the same. We consider . The estimator for the common variance is now
and the test statistic is
The distribution of this test statistic is similar to the distribution of the test statistic for the one-sample case: under the numerator is , is the sum of two independent and variables and thus is by definition A.2, and numerator and denominator are independent of each other by corollary 10.10, applied to each of the two samples, and the independence of the samples. Thus, the test can be used to test against , rejecting if , and the one-sided variants of the test use .
The third test is the -test for two variances. In the same two-sample setting, but now with variances and which are allowed to be different, the ratio is the ratio of two independent chi-squared variables, divided by the degrees of freedom. Thus, it is distributed by definition A.4. Under the unknown variances cancel and the statistic is distributed. The two-sided test rejects if does not fall into the interval , where is the -quantile of . Normally, tables only give the upper quantiles and the lower quantiles can be found by using . This test is also used to check the assumption of equal variances in (29.3).
29.5 Which Test?
Table 29.1 summarises the six tests. The choice between them is determined by the following three questions: Is the parameter in question a mean or a variance? Is there one sample or two? If two samples, are the observations paired or independent? All six tests assume that the data are normally distributed. In lecture 30 we will use simulation to study what happens when this assumption is violated.
| Question | Statistic | Null distribution | |
| one mean, known | |||
| one mean, unknown | |||
| paired differences | |||
| one variance, unknown | |||
| two means, common | |||
| two variances |
-
•
Each of the standard tests for normal data is based on the pivot from lecture 25, evaluated for the null value, and compared to a quantile of the known null distribution. The test is the inverse of the corresponding confidence interval.
-
•
The one-sample -test uses under and rejects for (one-sided) or (two-sided). It is the likelihood-ratio test with an exact critical value.
-
•
The chi-squared test for a variance uses under . A two-sided test requires both lower and upper quantiles, since the distribution is not symmetric.
-
•
The paired -test is a one-sample -test for the differences. Similarly, the two-sample -test with pooled variance and the -test for two variances follow this pattern, using the and distributions.
-
•
The table lists only a few quantiles and often the -value is bracketed between two tabulated levels.
Eight independent measurements of a quantity with nominal value are given by
Assume that the measurements are i.i.d. , where both parameters are unknown. From the -table we find , , and .
-
1.
Determine , and the observed value of the test statistic from (29.1) for .
-
2.
For , test at levels and . Bracket the -value.
-
3.
For , test at the same two levels and bracket the -value.
-
4.
Calculate the -interval for from example 25.4 and discuss how it relates to your answer in the previous part.
A sample of i.i.d. observations has sample variance , and the process generating the sample is known to have variance at most . From the chi-squared table we find , , , and .
-
1.
Test against at levels and . Bracket the -value.
-
2.
Test against at level .
-
3.
Show that the power function of the one-sided test from proposition 29.3 is given by where , and that this power function is increasing in . Deduce that the test has size for the composite null hypothesis .
-
4.
Using the power function, explain how it can happen that the one-sided test from part (a) and the two-sided test from part (b) can give different results at the same level.
The following data give the reaction times of six subjects before and after a training period.
| subject | 1 | 2 | 3 | 4 | 5 | 6 |
| before () | 72 | 65 | 80 | 58 | 69 | 75 |
| after () | 68 | 66 | 74 | 55 | 63 | 70 |
We want to know whether the training reduced the mean reaction time. From the -table we find , , and .
-
1.
Write down the model and hypotheses for the paired -test and carry out the test at levels and .
-
2.
Determine a confidence interval for the mean reduction in reaction time.
-
3.
Your colleague, not knowing that the two rows in the data frame give paired data, performs a two-sample -test (29.3). The sample variances of the two rows are and , respectively. What is the test statistic and the decision of your colleague at level ? Justify your answer by showing that the paired analysis is appropriate, and that the two analyses are different.
In example 25.4 we have observed a sample of values, where and , and we have found the -interval for . In example 25.5, for the same data, we have found the interval for . In exercise 25.4 we have seen that the hypothesis was rejected at level , using theorem 25.6 alone. Now use the values , , and .
-
1.
Perform a two-sided -test for and state the decision. Repeat the test for and check that the result is consistent with the interval.
-
2.
For , perform a two-sided test at level for and , either from the interval alone or by computing the test statistic (29.2).
-
3.
Perform a one-sided test for against at level . Why cannot the two-sided interval be used to answer this test? Which confidence set does this correspond to?