Lecture 23 Likelihood-Ratio Tests
Lecture 22 ended with a gap: the Neyman–Pearson lemma does not consider composite null hypotheses or two-sided alternatives, but these are the most common situations in statistical tests. The purpose of this lecture is to describe a generalisation of the likelihood-ratio test, which can be used for such cases. The generalised likelihood-ratio test, introduced in this section, is based on the same principles as the Neyman–Pearson ratio, but is modified to account for the more general case. In the second part of this lecture, we will see that the large-sample behaviour of the generalised likelihood-ratio test, described by Wilks’ theorem, is based on the asymptotic theory from lecture 17. The generalised likelihood-ratio test is the test computed by most statistical software packages when a likelihood-ratio test is requested.
23.1 The Generalised Likelihood Ratio
As in lecture 22, we write for the likelihood of the data and for the log-likelihood, where the parameter space is now allowed to be a subset of . We consider the following testing problem:
for a subset , where both hypotheses can be composite. Since the alternative is the complement of the null hypothesis within , the null hypothesis is a restriction of the full model and we have nested hypotheses.
If the hypothesis is composite, the likelihood principle suggests representing it in the Neyman–Pearson comparison of the two likelihoods by the parameter value which gives the best fit to the data, i.e. by the maximum of the likelihood over the hypothesis.
The generalised likelihood ratio for testing against is given by
The generalised likelihood-ratio test at level rejects , if , where the constant is chosen such that . The test can equivalently be written as rejecting for large values of the statistic
The denominator of is the likelihood at the maximum likelihood estimator from lecture 5, the numerator is the likelihood at the restricted maximum likelihood estimator , which is the best value within ; if is simple, we have and the numerator is just . Since we always have : values close to one mean that the restriction to costs little in terms of fit, values close to zero mean that is a much worse fit for the data than the full model is, and this is the evidence on which we base the rejection. The factor in is the scaling which, as we will see in section 23.3, gives the limiting distribution of its standard form. In terms of the log-likelihood we have
which is twice the amount the maximised log-likelihood decreases when the null hypothesis is imposed.
The test is an extension of the Neyman–Pearson test, not a replacement. For example, for a two-point parameter space where , and using the notation for the likelihood ratio from lecture 22, we have
and for the event is the same as the event ; thus, for this case, the test is the most powerful test from theorem 22.2 with . For composite hypotheses no optimality is guaranteed and in fact proposition 22.5 shows that for the two-sided normal problem no such test can exist.
Two more properties follow directly from the definition. First, if is a sufficient statistic, then we can use the factorisation theorem 8.2 to write , where the factor can be cancelled between numerator and denominator. This shows that can be written as a function of the data which depends on the data only through . Secondly, if is a one-to-one reparametrisation, the suprema in the definition of do not change and thus and do not depend on the choice of parametrisation of the model. This is the same invariance as in theorem 5.6, and exercise 23.3 verifies it in an example.
23.2 The Two-Sided Test for the Mean of a Normal Distribution
We will now consider the question which was left open at the end of lecture 22, i.e. we will consider the two-sided alternative for the mean of a normal sample with known variance. This example will carry over as an example for all the distributions considered later, since for this case the distribution of under can be found exactly.
Let be i.i.d. with known and consider against the alternative . Thus we have and . From example 5.4 we know that the log-likelihood
attains its maximum for at . Since the null hypothesis is simple, the restricted maximiser is and using formula (23.1) we find
Using the identity (2.1) from lecture 2, where is substituted for , we find that the difference of the two sums of squares is . Thus we have
The test rejects for large values of , i.e. for large values of . Thus we have identified the two-sided -test from example 20.7: the likelihood principle allowed us to derive the test which in lecture 22 was derived by common sense.
Under , the statistic is standard normally distributed by proposition A.3 and thus, by definition A.2, its square satisfies for every sample size . Let be the -quantile of the distribution. Then the level test rejects if . Since , the events and coincide and thus we have ; from the tables we find . For , , and observed mean , for the data from lecture 20, we find , which is larger than and thus is rejected at level . The -value is , which is the same value as obtained for the two-sided test in section 20.5.
A familiar test has been recovered from the general recipe, and the next section shows that the chi-squared distribution of is the general case for large samples, even in models where normality is not exactly satisfied.
23.3 Wilks’ Theorem
In example 23.2 the distribution of under was known, because is exactly normal; in most models this will not be the case and Wilks’ theorem provides the large-sample approximation which allows the test to be used practically.
The theorem is stated for hypotheses which fix some coordinates of the parameter. Let be open, write , and let the null hypothesis fix the first coordinates at given values,
where . For the null hypothesis is simple, and for the free coordinates are nuisance parameters, unknown parameters about which the hypothesis says nothing. The number of constraints is the difference between the dimensions of and .
Let be i.i.d. from a model with parameter which satisfies the regularity conditions of definition 16.1, extended to a -dimensional parameter as at the end of lecture 17, and suppose that, for every sample where all coordinates are in the support of the model, the maximum likelihood estimator is the unique root of the likelihood equations, both over and over . Let be of the form (23.3) and let be the generalised likelihood-ratio statistic for a sample of size . Then for every we have
under .
Whatever the model, the only feature of the problem which enters the limit is the number of constraints. In example 23.2 we have and there the statement holds exactly for every . We do not prove the theorem: together with the vector case of theorem 17.1, on which its general form rests, it is one of the few results of this module which we state without proof. The following argument shows where the chi-squared distribution comes from.
Consider the case of a one-dimensional parameter and a simple null hypothesis , so that . We will expand the log-likelihood around the maximum likelihood estimator (MLE) instead of around as in lecture 17: using Taylor’s theorem with the Lagrange form of the remainder we find
for some between and . Since solves the likelihood equation, the first derivative at this point is zero and rearranging this gives
Now assume that is true. Then, by theorem 17.1, the second term is the square of a quantity which converges in distribution to , and by the continuous mapping theorem, theorem A.17, this shows that the second term converges in distribution to where is standard normally distributed. The first term is the term from the proof of theorem 17.1 with instead of and since is between and the consistent estimator , the bound (R4) on the third derivative shows that this term still converges in probability to . Finally, by Slutsky’s theorem, theorem A.16, we find
by definition A.2. Since the Fisher information cancels, the limit is independent of the details of the model.
For a simple null hypothesis in one dimension, this argument is close to a complete proof. The proof for the general case is more challenging: the expansion will be in variables and will use the vector version of theorem 17.1, and the quadratic form which replaces will consist of independent, squared standard normal variables, one for each constraint. This is the origin of the number of degrees of freedom in the argument.
23.4 Using the Theorem
Wilks’ theorem turns the generalised likelihood-ratio test into a recipe: compute the statistic from the two maximised log-likelihoods, found numerically if need be as in lecture 18; count the degrees of freedom ; reject at level , if ; and report the approximate -value for the observed value , which for equals . Only the counting of degrees of freedom requires some care. For example, in the case of the mean of a normal sample with unknown variance, and , we have and , since the variance is unrestricted under both hypotheses. From exercise 23.2 we know that for a multinomial distribution with categories and a fully specified null hypothesis.
The hypothesis (23.3) fixes some of the coordinates, but since is invariant under reparametrisation, the theorem applies to all smooth, nested hypotheses of dimension . For example, the line in the plane of two normal means can be transformed to with free by a change of coordinates. Thus, the rule can be applied directly.
The theorem is a limit statement and there are three cases where problems can occur when applying the theorem to a fixed . First, the approximation may be poor for small sample size: for the mean of a normal sample with unknown variance, exercise 23.1 shows that the level test based on rejects too often at . Lecture 24 watches the approximation improve as increases, by simulation. Whenever the exact distribution of or of an equivalent test statistic is known, as in the situation of the exercise, this distribution should be used. Lecture 29 takes this approach, using the exact distribution, for the standard tests. Secondly, if the regularity conditions are violated, for example if the support of the distribution depends on the parameter or if the null value coincides with the boundary of the parameter space, the expansion (23.4) will break down and the limit will not be . Finally, the theorem only describes under . The theorem gives the size of the test, but the power of the test must be determined separately.
Software output often lists two relatives of the likelihood-ratio test: the Wald test, using as in corollary 17.3, and the score test, using the squared score . Both tests have the same chi-squared limit under as , but here we use the likelihood-ratio test throughout.
-
•
For nested hypotheses against , the generalised likelihood ratio is , and the generalised likelihood-ratio test rejects for small , i.e. for large .
-
•
For two simple hypotheses the test is the Neyman–Pearson test; in general it depends on the data only through a sufficient statistic and is unchanged by reparametrisation.
-
•
For a normal mean with known variance and , the statistic is , which is exactly under and the test is the two-sided -test.
-
•
Wilks’ theorem: under regularity conditions, if fixes of the parameters, then under . The theorem is given without proof, but from the arguments for , using Taylor expansion of around and theorem 17.1, it is clear why the limit is .
-
•
In use: compute from the two maximised log-likelihoods, count , reject if . The approximation can be poor for small , the test fails if the regularity conditions are violated, and no information about power is available.
Let be i.i.d. , where both parameters are unknown. Consider the hypothesis against the alternative , where is a nuisance parameter. Let , and .
-
1.
Following example 5.4, show that the maximised log-likelihood function over is given by
-
2.
Show that, restricting ourselves to the region , the maximum of the likelihood function is attained for and find expressions for in terms of and using the identity (2.1).
-
3.
Deduce that
and that the generalised likelihood-ratio test rejects for large values of .
-
4.
State the approximate distribution of under given by theorem 23.3 and justify your answer for the degrees of freedom. Recall that, under , the statistic follows the distribution by definition A.4 and corollary 10.10. For , determine the value of above which the approximate test with critical value rejects and compare your result with the -quantile of the distribution. Comment upon your result.
An experiment with possible outcomes is repeated times independently. The number of times each outcome occurs is given by for outcome . Thus, is multinomially distributed with probability weights
where satisfies and . We test for a given vector against the alternative .
-
1.
Show that the log-likelihood function is up to an additive constant and use the Lagrange multiplier method, with constraint , to show that the maximum of the likelihood function is at .
-
2.
Show that
where .
-
3.
Why does theorem 23.3 state that under ?
-
4.
A die is thrown times and the six faces appear , , , , and times, respectively. Test the hypothesis that the die is fair at level , using .
-
5.
Let so that . Show that, by expanding the logarithm to second order in , we find
This is Pearson’s chi-squared statistic. Compute the value for the die in the previous part of the question.
Let be i.i.d. with Bernoulli distribution, i.e. with success probability . Furthermore, let and and consider the hypothesis against the alternative .
-
1.
Show that
where we use the convention .
-
2.
Approximately determine the distribution of under for large and use the resulting distribution to derive a test at level .
-
3.
In trials, we observed successes. Test the hypothesis at level and determine the approximate -value.
-
4.
In exercise 17.1, the model was reparametrised using the log-odds . Show that the generalised likelihood-ratio statistic for testing coincides with the one in part (a) for , the value of the null proportion.