Lecture 22 The Neyman–Pearson Lemma
In lecture 20 it was clear that for the mean of a normal sample we could not do better than to reject for large values of . This lecture proves that this is indeed correct: for the most basic testing problem, where one distribution is being tested against a different distribution, the Neyman–Pearson lemma allows us to identify the most powerful test. This is in analogy to the role of the Cramer–Rao bound from lecture 13 for estimators. The same test is optimal for all one-sided alternatives, but there is no single test which is optimal for a two-sided alternative. In fact, in lecture 23 we will see the required tool for this case.
22.1 Simple Hypotheses and Most Powerful Tests
We consider the problem of testing a simple null hypothesis against a simple alternative , where are two fixed points of the parameter space, for data with joint density or probability weights as in lecture 8. Under either hypothesis the distribution of the data is completely specified; section 22.4 shows that the solution of this problem also settles more realistic ones.
We will compare different tests, so we write for the power function of definition 20.4 of the test with critical region . For simple hypotheses, the size from definition 20.5 and the power are the single numbers and . Among all tests with level , i.e. with , we want to find the test which makes type II errors least often.
Let . A test with critical region is said to be most powerful at level for testing against , if and
for all critical regions with .
As with the uniformly minimum variance unbiased estimator from definition 10.3, the definition does not give a method to find such a test or even whether such a test exists; we will answer both of these questions in the next section.
22.2 The Lemma
If is much larger than , the data favour the alternative, and we should reject. The natural test statistic is the likelihood ratio
and the natural test rejects if exceeds some threshold . The lemma states that no test with the same level can have more power. We write the inequality instead of so that points with do not require separate treatment and so that the boundary of the critical region can include any points where equality holds, as required for discrete models and in exercise 22.3.
Let and let be a critical region with
where is the size of the critical region. Then the following statements hold.
-
1.
The test with critical region is the most powerful test at level for the hypothesis against .
-
2.
If is another critical region, which is most powerful at level for the same problem, then and differ at most on the boundary set , except for a set of observations with probability zero for every .
Let be a critical region with and let and be the indicator functions for the two regions. Then the proof is based on the observation that
If , the first factor is either or and the second factor is non-negative by the choice of . If , the first factor is either or and the second factor is zero or negative by the choice of .
Integrating (22.1) over all , with a sum in place of the integral for a discrete model, and using for each of the four terms, we find
Since and , the subtracted term is non-negative and thus we have
and this is the first statement.
For the second statement assume that is also most powerful at level . Then by definition 22.1, and by the first statement, so that the two powers are equal and the first bracket in (22.2) vanishes. The right-hand side of (22.2) is then , and together with the inequality on the left this gives . Now holds exactly for those which lie in one of the regions and but not in the other and at which , since at such points both factors of are non-zero and, by the first paragraph, of the same sign. For a discrete model a sum of non-negative terms is zero only if every term is zero, and thus this set is empty. For a model with densities, a non-negative function with integral zero vanishes outside a set of measure zero, and a set of measure zero has probability zero under every density . This completes the proof. ∎
Two comments help in using the lemma. First, the constant is rarely computed: whenever is a strictly monotone function of a simpler statistic , we rewrite as or and choose from the distribution of under to give size . Second, the lemma shows that the test is most powerful at the level equal to its own size, whatever value that size takes. For models with densities, usually has a continuous distribution, and thus in theory any size can be achieved and the boundary set in the second statement has probability zero. Together these results show that the second statement implies that the most powerful test is essentially unique. In addition, it can be shown (for example in exercise 22.4) that the most powerful test can be chosen to depend on the data only via a sufficient statistic.
For discrete models the size of tests can only take certain values as changes and it is in general not possible to obtain a test with a given size . The usual solution in textbooks, to randomise the test on the boundary set to achieve the required size, is not taken here. Instead, we consider the test with the largest possible size which is less than or equal to . This test, by the result of exercise 22.1, will be the most powerful test for the smaller level.
22.3 The Normal Mean
We now apply the lemma to the normal-mean example from lecture 20: the test constructed there using common sense is the most powerful one.
Let be i.i.d. where is known, and consider the hypothesis against the alternative , where . The joint density of the observations is . In the ratio of two densities, the constant factors cancel. After expanding the squares in the definition of the test statistic, we find
and thus
Since , the right-hand side is a strictly increasing function of and the inequality is equivalent to where . Thus, the critical region is of the form and we can choose the value of by considering the size of the test: under we have by proposition A.3 and thus
if and only if
where is the -quantile of the standard normal distribution. By theorem 22.2 the test which rejects for is most powerful at level and, since the boundary set has probability zero, this is the only test up to sets of probability zero. This is the test from example 20.6. The power of this test for the alternative is
For , , and we have and thus , and the power against is , the value we found in lecture 20. The lemma allows us to turn this number into a statement about all possible procedures: no test of level , based on the given observations, can be better.
22.4 One-Sided Alternatives
The threshold in example 22.3 does not involve : the alternative entered the computation only through the sign of , and thus every leads to the same critical region. This has a consequence for the composite problem of testing against , the problem of interest in lecture 20.
Consider testing against . Assume that is a critical region with which, for every , satisfies the condition from theorem 22.2 for testing against , for some constant which may depend on . Then, for every and every critical region with , we have .
Let and let be a critical region with . Then, since the null hypothesis is simple, is also a test of level for the simple hypothesis against and by assumption satisfies the condition from theorem 22.2 for this problem. Thus, by the first statement of the theorem, we have and since was arbitrary this completes the proof. ∎
A test which has this property is most powerful against every alternative in simultaneously and is said to be uniformly most powerful at level . We only use this name as a description and do not consider the general theory of such tests.
For the normal mean, example 22.3 shows that the critical region is the Neyman–Pearson region for every , and thus the one-sided test of example 20.6 is uniformly most powerful against , with power function
For alternatives the same computation with the inequality reversed gives the region , uniformly most powerful against , with power function
22.5 Two-Sided Alternatives
We now consider the alternative . The two one-sided tests disagree: one rejects for large values of , the other for small values, and each one is poor on the other side of , since for we have . Thus, no single test can agree with both tests simultaneously, and the uniqueness statement in the lemma implies that this argument constitutes a proof.
Let be i.i.d. with known and let . Then there is no critical region with such that for all and any critical region with .
Suppose that were such a region, and fix some . Then is most powerful at level for testing against , and so is the region of example 22.3. The boundary set has probability zero, and thus by the second statement of theorem 22.2 the regions and differ only on a set which has probability zero for every . Consequently the two tests have the same power function, for all . Now take any . From the formulas of section 22.4 we get
since is strictly increasing and . The left-sided test has size and thus is an admissible competitor , and it has strictly larger power at than . This contradicts the assumed property of , which completes the proof. ∎
The argument is not particular to the normal distribution: whenever the one-sided most powerful regions disagree, a uniformly most powerful two-sided test would have to coincide with both, which is impossible, and so we have to settle for a compromise. The two-sided test in example 20.7 is such a compromise: the power of this test is less than for , but it is strictly larger than for all . In lecture 23 we will see that this test is the result of applying the likelihood principle to a composite alternative. The generalised likelihood-ratio test obtained in this way is a general-purpose tool which can be used when no uniformly most powerful test exists.
-
•
A test is most powerful at level for a simple null hypothesis against a simple alternative, if it has level and no other test of level has larger power.
-
•
The Neyman–Pearson lemma: The test which rejects when , where is chosen so that the size is , is most powerful at level . Any other most powerful test coincides with this test outside the boundary set. The proof involves integration of .
-
•
In practice, can be rewritten as for a simple statistic and can be found by considering the distribution of under ; for the normal mean this gives .
-
•
If the region of acceptance does not depend on the alternative , the same test will be most powerful against all : the test is then uniformly most powerful for the one-sided problem.
-
•
For the two-sided alternative, no uniformly most powerful test can exist, since the two one-sided tests disagree; instead, the likelihood-ratio tests from lecture 23 can be used.
Let be i.i.d. Poisson with mean and let , Poisson with mean . We test against where .
-
1.
Show that the likelihood ratio is a strictly increasing function of and that the most powerful test rejects when . Why does not depend on ? What are the consequences for testing against the composite alternative ?
-
2.
Let , and . Determine the smallest integer such that the test has size at most . Compute the actual size of the test for this value. You may use and for a Poisson random variable with mean .
-
3.
Determine the power of the test at , using for a Poisson random variable with mean .
-
4.
Why cannot a test of the form have size exactly ? At which level is the test from part (b) most powerful?
Let be i.i.d. Bernoulli with success probability , and let . In exercise 20.2 we tested against by rejecting for large values of .
-
1.
For a fixed alternative , show that the likelihood ratio is given by
and that it is strictly increasing in . Deduce that the most powerful test rejects when , and that the test of exercise 20.2 is uniformly most powerful against .
-
2.
Let . Using and , find the critical region of the largest size not exceeding , and compute its power at using . Compare the size and the power with those found for in exercise 20.2.
-
3.
For an alternative , find the most powerful critical region, and explain why the two one-sided problems lead to different tests.
Let be i.i.d. uniformly distributed on the set , with joint density for , where , and let . From example 8.10 we know that for . We want to test against , where , at significance level . In exercise 20.3 we have considered tests based on for this model, and in the following we will see that these tests are optimal.
-
1.
Let . Determine the sets and in terms of .
-
2.
Show that the critical region for has size and satisfies the condition from theorem 22.2 for this . Conclude that this is the most powerful test at level . Compute the power of the test and explain why the test is uniformly most powerful against .
-
3.
Show that the region also has size and the same power as . Explain, using the second statement of theorem 22.2, why this does not contradict the theorem, even though the model has densities.
Let be a sufficient statistic for the model and let be two parameter values. Using the factorisation theorem 8.2, show that the likelihood ratio depends on the data only through , for all where the denominator is positive.