Lecture 23 Likelihood-Ratio Tests

Lecture 22 ended with a gap: the Neyman–Pearson lemma does not consider composite null hypotheses or two-sided alternatives, but these are the most common situations in statistical tests. The purpose of this lecture is to describe a generalisation of the likelihood-ratio test, which can be used for such cases. The generalised likelihood-ratio test, introduced in this section, is based on the same principles as the Neyman–Pearson ratio, but is modified to account for the more general case. In the second part of this lecture, we will see that the large-sample behaviour of the generalised likelihood-ratio test, described by Wilks’ theorem, is based on the asymptotic theory from lecture 17. The generalised likelihood-ratio test is the test computed by most statistical software packages when a likelihood-ratio test is requested.

23.1 The Generalised Likelihood Ratio

As in lecture 22, we write L⁢(θ)L(\theta) for the likelihood of the data X=(X1,…,Xn)X=(X_{1},\dots,X_{n}) and ℓ⁢(θ)=log⁡L⁢(θ)\ell(\theta)=\log L(\theta) for the log-likelihood, where the parameter space Θ\Theta is now allowed to be a subset of ℝk\mathbb{R}^{k}. We consider the following testing problem:

H0:θ∈Θ0againstH1:θ∈Θ∖Θ0H_{0}\colon\theta\in\Theta_{0}\qquad\text{against}\qquad H_{1}\colon\theta\in% \Theta\setminus\Theta_{0}

for a subset Θ0⊆Θ\Theta_{0}\subseteq\Theta, where both hypotheses can be composite. Since the alternative is the complement of the null hypothesis within Θ\Theta, the null hypothesis is a restriction of the full model and we have nested hypotheses.

If the hypothesis is composite, the likelihood principle suggests representing it in the Neyman–Pearson comparison of the two likelihoods by the parameter value which gives the best fit to the data, i.e. by the maximum of the likelihood over the hypothesis.

Definition 23.1.

The generalised likelihood ratio for testing H0:θ∈Θ0H_{0}\colon\theta\in\Theta_{0} against H1:θ∈Θ∖Θ0H_{1}\colon\theta\in\Theta\setminus\Theta_{0} is given by

Λ⁢(x)=supθ∈Θ0L⁢(θ)supθ∈ΘL⁢(θ).\Lambda(x)=\frac{\sup_{\theta\in\Theta_{0}}L(\theta)}{\sup_{\theta\in\Theta}L(% \theta)}.

The generalised likelihood-ratio test at level α\alpha rejects H0H_{0}, if Λ⁢(X)≤c\Lambda(X)\leq c, where the constant cc is chosen such that supθ∈Θ0ℙθ⁢(Λ⁢(X)≤c)≤α\sup_{\theta\in\Theta_{0}}\mathbb{P}_{\theta}\bigl{(}\Lambda(X)\leq c\bigr{)}\leq\alpha. The test can equivalently be written as rejecting H0H_{0} for large values of the statistic

W=−2⁢log⁡Λ⁢(X).W=-2\log\Lambda(X).

The denominator of Λ\Lambda is the likelihood at the maximum likelihood estimator θ^\hat{\theta} from lecture 5, the numerator is the likelihood at the restricted maximum likelihood estimator θ^0\hat{\theta}_{0}, which is the best value within H0H_{0}; if H0H_{0} is simple, we have Θ0={θ0}\Theta_{0}=\{\theta_{0}\} and the numerator is just L⁢(θ0)L(\theta_{0}). Since Θ0⊆Θ\Theta_{0}\subseteq\Theta we always have 0≤Λ≤10\leq\Lambda\leq 1: values close to one mean that the restriction to Θ0\Theta_{0} costs little in terms of fit, values close to zero mean that H0H_{0} is a much worse fit for the data than the full model is, and this is the evidence on which we base the rejection. The factor −2-2 in WW is the scaling which, as we will see in section 23.3, gives the limiting distribution of WW its standard form. In terms of the log-likelihood we have

equation (23.1) (23.1)
W=2⁢(ℓ⁢(θ^)−ℓ⁢(θ^0)),W=2\bigl{(}\ell(\hat{\theta})-\ell(\hat{\theta}_{0})\bigr{)},

which is twice the amount the maximised log-likelihood decreases when the null hypothesis is imposed.

The test is an extension of the Neyman–Pearson test, not a replacement. For example, for a two-point parameter space Θ={θ0,θ1}\Theta=\{\theta_{0},\theta_{1}\} where Θ0={θ0}\Theta_{0}=\{\theta_{0}\}, and using the notation R=L⁢(θ1)/L⁢(θ0)R=L(\theta_{1})/L(\theta_{0}) for the likelihood ratio from lecture 22, we have

Λ=L⁢(θ0)max⁡(L⁢(θ0),L⁢(θ1))=min⁡(1,1R),\Lambda=\frac{L(\theta_{0})}{\max\bigl{(}L(\theta_{0}),L(\theta_{1})\bigr{)}}=% \min\Bigl{(}1,\frac{1}{R}\Bigr{)},

and for c<1c<1 the event Λ≤c\Lambda\leq c is the same as the event R≥1/cR\geq 1/c; thus, for this case, the test is the most powerful test from theorem 22.2 with k=1/ck=1/c. For composite hypotheses no optimality is guaranteed and in fact proposition 22.5 shows that for the two-sided normal problem no such test can exist.

Two more properties follow directly from the definition. First, if TT is a sufficient statistic, then we can use the factorisation theorem 8.2 to write L⁢(θ)=g⁢(T⁢(x);θ)⁢h⁢(x)L(\theta)=g\bigl{(}T(x);\theta\bigr{)}h(x), where the factor h⁢(x)h(x) can be cancelled between numerator and denominator. This shows that Λ\Lambda can be written as a function of the data which depends on the data only through TT. Secondly, if ψ=g⁢(θ)\psi=g(\theta) is a one-to-one reparametrisation, the suprema in the definition of Λ\Lambda do not change and thus Λ\Lambda and WW do not depend on the choice of parametrisation of the model. This is the same invariance as in theorem 5.6, and exercise 23.3 verifies it in an example.

23.2 The Two-Sided Test for the Mean of a Normal Distribution

We will now consider the question which was left open at the end of lecture 22, i.e. we will consider the two-sided alternative for the mean of a normal sample with known variance. This example will carry over as an example for all the distributions considered later, since for this case the distribution of WW under H0H_{0} can be found exactly.

Example 23.2.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. N⁢(θ,σ2)N(\theta,\sigma^{2}) with known σ2\sigma^{2} and consider H0:θ=θ0H_{0}\colon\theta=\theta_{0} against the alternative H1:θ≠θ0H_{1}\colon\theta\neq\theta_{0}. Thus we have Θ=ℝ\Theta=\mathbb{R} and Θ0={θ0}\Theta_{0}=\{\theta_{0}\}. From example 5.4 we know that the log-likelihood

ℓ⁢(θ)=−n2⁢log⁡(2⁢π⁢σ2)−12⁢σ2⁢∑i=1n(Xi−θ)2\ell(\theta)=-\frac{n}{2}\log(2\pi\sigma^{2})-\frac{1}{2\sigma^{2}}\sum_{i=1}^% {n}(X_{i}-\theta)^{2}

attains its maximum for ℝ\mathbb{R} at θ^=X¯\hat{\theta}=\bar{X}. Since the null hypothesis is simple, the restricted maximiser is θ0\theta_{0} and using formula (23.1) we find

W=2⁢(ℓ⁢(X¯)−ℓ⁢(θ0))=1σ2⁢(∑i=1n(Xi−θ0)2−∑i=1n(Xi−X¯)2).W=2\bigl{(}\ell(\bar{X})-\ell(\theta_{0})\bigr{)}=\frac{1}{\sigma^{2}}\Bigl{(}% \sum_{i=1}^{n}(X_{i}-\theta_{0})^{2}-\sum_{i=1}^{n}(X_{i}-\bar{X})^{2}\Bigr{)}.

Using the identity (2.1) from lecture 2, where θ0\theta_{0} is substituted for μ\mu, we find that the difference of the two sums of squares is n⁢(X¯−θ0)2n(\bar{X}-\theta_{0})^{2}. Thus we have

equation (23.2) (23.2)
W=n⁢(X¯−θ0)2σ2=Z2,where ⁢Z=n⁢(X¯−θ0)σ.W=\frac{n(\bar{X}-\theta_{0})^{2}}{\sigma^{2}}=Z^{2},\qquad\text{where }Z=% \frac{\sqrt{n}\,(\bar{X}-\theta_{0})}{\sigma}.

The test rejects for large values of WW, i.e. for large values of |Z||Z|. Thus we have identified the two-sided zz-test from example 20.7: the likelihood principle allowed us to derive the test which in lecture 22 was derived by common sense.

Under ℙθ0\mathbb{P}_{\theta_{0}}, the statistic ZZ is standard normally distributed by proposition A.3 and thus, by definition A.2, its square W=Z2W=Z^{2} satisfies W∼χ12W\sim\chi^{2}_{1} for every sample size nn. Let χr,p2\chi^{2}_{r,p} be the pp-quantile of the χr2\chi^{2}_{r} distribution. Then the level α\alpha test rejects if W≥χ1,1−α2W\geq\chi^{2}_{1,1-\alpha}. Since W=Z2W=Z^{2}, the events W≥χ1,1−α2W\geq\chi^{2}_{1,1-\alpha} and |Z|≥z1−α/2|Z|\geq z_{1-\alpha/2} coincide and thus we have χ1,1−α2=z1−α/22\chi^{2}_{1,1-\alpha}=z_{1-\alpha/2}^{2}; from the tables we find χ1,0.952=3.841=1.962\chi^{2}_{1,0.95}=3.841=1.96^{2}. For n=25n=25, σ=2\sigma=2, θ0=10\theta_{0}=10 and observed mean x¯=10.9\bar{x}=10.9, for the data from lecture 20, we find w=25⋅0.81/4=5.06w=25\cdot 0.81/4=5.06, which is larger than 3.8413.841 and thus H0H_{0} is rejected at level 0.050.05. The pp-value is ℙ⁢(χ12≥5.06)=2⁢(1−Φ⁢(5.06))=2⁢(1−Φ⁢(2.25))=0.0244\mathbb{P}(\chi^{2}_{1}\geq 5.06)=2\bigl{(}1-\Phi(\sqrt{5.06})\bigr{)}=2\bigl{% (}1-\Phi(2.25)\bigr{)}=0.0244, which is the same value as obtained for the two-sided test in section 20.5.

A familiar test has been recovered from the general recipe, and the next section shows that the chi-squared distribution of WW is the general case for large samples, even in models where normality is not exactly satisfied.

23.3 Wilks’ Theorem

In example 23.2 the distribution of WW under H0H_{0} was known, because X¯\bar{X} is exactly normal; in most models this will not be the case and Wilks’ theorem provides the large-sample approximation which allows the test to be used practically.

The theorem is stated for hypotheses which fix some coordinates of the parameter. Let Θ⊆ℝk\Theta\subseteq\mathbb{R}^{k} be open, write θ=(θ1,…,θk)\theta=(\theta_{1},\dots,\theta_{k}), and let the null hypothesis fix the first rr coordinates at given values,

equation (23.3) (23.3)
Θ0={θ∈Θ|θj=θj0⁢ for all ⁢j≤r},\Theta_{0}=\bigl{\{}\,\theta\in\Theta\mathrel{\big{|}}\theta_{j}=\theta_{j}^{0% }\text{ for all }j\leq r\,\bigr{\}},

where 1≤r≤k1\leq r\leq k. For r=kr=k the null hypothesis is simple, and for r<kr<k the free coordinates θr+1,…,θk\theta_{r+1},\dots,\theta_{k} are nuisance parameters, unknown parameters about which the hypothesis says nothing. The number rr of constraints is the difference between the dimensions of Θ\Theta and Θ0\Theta_{0}.

Theorem 23.3 (Wilks).

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. from a model with parameter θ∈Θ⊆ℝk\theta\in\Theta\subseteq\mathbb{R}^{k} which satisfies the regularity conditions of definition 16.1, extended to a kk-dimensional parameter as at the end of lecture 17, and suppose that, for every sample where all coordinates are in the support of the model, the maximum likelihood estimator is the unique root of the likelihood equations, both over Θ\Theta and over Θ0\Theta_{0}. Let Θ0\Theta_{0} be of the form (23.3) and let Wn=−2⁢log⁡ΛnW_{n}=-2\log\Lambda_{n} be the generalised likelihood-ratio statistic for a sample of size nn. Then for every θ∈Θ0\theta\in\Theta_{0} we have

Wn→dχr2(n→∞)W_{n}\xrightarrow{\ \mathrm{d}\ }\chi^{2}_{r}\qquad(n\to\infty)

under ℙθ\mathbb{P}_{\theta}.

Whatever the model, the only feature of the problem which enters the limit is the number rr of constraints. In example 23.2 we have k=r=1k=r=1 and there the statement holds exactly for every nn. We do not prove the theorem: together with the vector case of theorem 17.1, on which its general form rests, it is one of the few results of this module which we state without proof. The following argument shows where the chi-squared distribution comes from.

Consider the case k=r=1k=r=1 of a one-dimensional parameter and a simple null hypothesis H0:θ=θ0H_{0}\colon\theta=\theta_{0}, so that Wn=2⁢(ℓ⁢(θ^n)−ℓ⁢(θ0))W_{n}=2\bigl{(}\ell(\hat{\theta}_{n})-\ell(\theta_{0})\bigr{)}. We will expand the log-likelihood around the maximum likelihood estimator (MLE) instead of around θ0\theta_{0} as in lecture 17: using Taylor’s theorem with the Lagrange form of the remainder we find

ℓ⁢(θ0)=ℓ⁢(θ^n)+ℓ′⁢(θ^n)⁢(θ0−θ^n)+12⁢ℓ′′⁢(θn∗)⁢(θ0−θ^n)2\ell(\theta_{0})=\ell(\hat{\theta}_{n})+\ell^{\prime}(\hat{\theta}_{n})(\theta% _{0}-\hat{\theta}_{n})+\tfrac{1}{2}\,\ell^{\prime\prime}(\theta_{n}^{*})(% \theta_{0}-\hat{\theta}_{n})^{2}

for some θn∗\theta_{n}^{*} between θ0\theta_{0} and θ^n\hat{\theta}_{n}. Since θ^n\hat{\theta}_{n} solves the likelihood equation, the first derivative at this point is zero and rearranging this gives

equation (23.4) (23.4)
Wn=−ℓ′′⁢(θn∗)⁢(θ^n−θ0)2=(−1n⁢ℓ′′⁢(θn∗))⁢(n⁢(θ^n−θ0))2.W_{n}=-\ell^{\prime\prime}(\theta_{n}^{*})\,(\hat{\theta}_{n}-\theta_{0})^{2}=% \Bigl{(}-\frac{1}{n}\,\ell^{\prime\prime}(\theta_{n}^{*})\Bigr{)}\Bigl{(}\sqrt% {n}\,(\hat{\theta}_{n}-\theta_{0})\Bigr{)}^{2}.

Now assume that H0H_{0} is true. Then, by theorem 17.1, the second term is the square of a quantity which converges in distribution to N⁢(0,1/ℐ︀⁢(θ0))N\bigl{(}0,1/\mathcal{I}(\theta_{0})\bigr{)}, and by the continuous mapping theorem, theorem A.17, this shows that the second term converges in distribution to Y2/ℐ︀⁢(θ0)Y^{2}/\mathcal{I}(\theta_{0}) where YY is standard normally distributed. The first term is the term BnB_{n} from the proof of theorem 17.1 with θn∗\theta_{n}^{*} instead of θ0\theta_{0} and since θn∗\theta_{n}^{*} is between θ0\theta_{0} and the consistent estimator θ^n\hat{\theta}_{n}, the bound (R4) on the third derivative shows that this term still converges in probability to ℐ︀⁢(θ0)\mathcal{I}(\theta_{0}). Finally, by Slutsky’s theorem, theorem A.16, we find

Wn→dℐ︀⁢(θ0)⋅Y2ℐ︀⁢(θ0)=Y2∼χ12,W_{n}\xrightarrow{\ \mathrm{d}\ }\mathcal{I}(\theta_{0})\cdot\frac{Y^{2}}{% \mathcal{I}(\theta_{0})}=Y^{2}\sim\chi^{2}_{1},

by definition A.2. Since the Fisher information cancels, the limit is independent of the details of the model.

For a simple null hypothesis in one dimension, this argument is close to a complete proof. The proof for the general case is more challenging: the expansion will be in kk variables and will use the vector version of theorem 17.1, and the quadratic form which replaces Y2Y^{2} will consist of rr independent, squared standard normal variables, one for each constraint. This is the origin of the number rr of degrees of freedom in the argument.

23.4 Using the Theorem

Wilks’ theorem turns the generalised likelihood-ratio test into a recipe: compute the statistic W=2⁢(ℓ⁢(θ^)−ℓ⁢(θ^0))W=2\bigl{(}\ell(\hat{\theta})-\ell(\hat{\theta}_{0})\bigr{)} from the two maximised log-likelihoods, found numerically if need be as in lecture 18; count the degrees of freedom r=dimΘ−dimΘ0r=\dim\Theta-\dim\Theta_{0}; reject H0H_{0} at level α\alpha, if W≥χr,1−α2W\geq\chi^{2}_{r,1-\alpha}; and report the approximate pp-value ℙ⁢(χr2≥w)\mathbb{P}(\chi^{2}_{r}\geq w) for the observed value ww, which for r=1r=1 equals 2⁢(1−Φ⁢(w))2\bigl{(}1-\Phi(\sqrt{w})\bigr{)}. Only the counting of degrees of freedom requires some care. For example, in the case of the mean of a normal sample with unknown variance, and H0:μ=μ0H_{0}\colon\mu=\mu_{0}, we have k=2k=2 and r=1r=1, since the variance is unrestricted under both hypotheses. From exercise 23.2 we know that r=m−1r=m-1 for a multinomial distribution with mm categories and a fully specified null hypothesis.

Remark.

The hypothesis (23.3) fixes some of the coordinates, but since WW is invariant under reparametrisation, the theorem applies to all smooth, nested hypotheses of dimension k−rk-r. For example, the line {μ1=μ2}\{\mu_{1}=\mu_{2}\} in the plane of two normal means can be transformed to θ1=μ1−μ2=0\theta_{1}=\mu_{1}-\mu_{2}=0 with θ2=μ2\theta_{2}=\mu_{2} free by a change of coordinates. Thus, the rule r=dimΘ−dimΘ0r=\dim\Theta-\dim\Theta_{0} can be applied directly.

The theorem is a limit statement and there are three cases where problems can occur when applying the theorem to a fixed nn. First, the approximation may be poor for small sample size: for the mean of a normal sample with unknown variance, exercise 23.1 shows that the level 0.050.05 test based on χ1,0.952\chi^{2}_{1,0.95} rejects too often at n=10n=10. Lecture 24 watches the approximation improve as nn increases, by simulation. Whenever the exact distribution of WW or of an equivalent test statistic is known, as in the situation of the exercise, this distribution should be used. Lecture 29 takes this approach, using the exact distribution, for the standard tests. Secondly, if the regularity conditions are violated, for example if the support of the distribution depends on the parameter or if the null value coincides with the boundary of the parameter space, the expansion (23.4) will break down and the limit will not be χr2\chi^{2}_{r}. Finally, the theorem only describes WW under H0H_{0}. The theorem gives the size of the test, but the power of the test must be determined separately.

Software output often lists two relatives of the likelihood-ratio test: the Wald test, using n⁢ℐ︀⁢(θ^)⁢(θ^−θ0)2n\mathcal{I}(\hat{\theta})(\hat{\theta}-\theta_{0})^{2} as in corollary 17.3, and the score test, using the squared score ℓ′⁢(θ0)2/(n⁢ℐ︀⁢(θ0))\ell^{\prime}(\theta_{0})^{2}/\bigl{(}n\mathcal{I}(\theta_{0})\bigr{)}. Both tests have the same chi-squared limit under H0H_{0} as WW, but here we use the likelihood-ratio test throughout.

Summary.
  • •
    ​

    For nested hypotheses H0:θ∈Θ0H_{0}\colon\theta\in\Theta_{0} against H1:θ∈Θ∖Θ0H_{1}\colon\theta\in\Theta\setminus\Theta_{0}, the generalised likelihood ratio is Λ=supΘ0L/supΘL∈[0,1]\Lambda=\sup_{\Theta_{0}}L/\sup_{\Theta}L\in[0,1], and the generalised likelihood-ratio test rejects for small Λ\Lambda, i.e. for large W=−2⁢log⁡Λ=2⁢(ℓ⁢(θ^)−ℓ⁢(θ^0))W=-2\log\Lambda=2(\ell(\hat{\theta})-\ell(\hat{\theta}_{0})).

  • •
    ​

    For two simple hypotheses the test is the Neyman–Pearson test; in general it depends on the data only through a sufficient statistic and is unchanged by reparametrisation.

  • •
    ​

    For a normal mean with known variance and H0:θ=θ0H_{0}\colon\theta=\theta_{0}, the statistic is W=n⁢(X¯−θ0)2/σ2=Z2W=n(\bar{X}-\theta_{0})^{2}/\sigma^{2}=Z^{2}, which is exactly χ12\chi^{2}_{1} under H0H_{0} and the test is the two-sided zz-test.

  • •
    ​

    Wilks’ theorem: under regularity conditions, if H0H_{0} fixes rr of the kk parameters, then W→dχr2W\xrightarrow{\ \mathrm{d}\ }\chi^{2}_{r} under H0H_{0}. The theorem is given without proof, but from the arguments for k=r=1k=r=1, using Taylor expansion of ℓ\ell around θ^n\hat{\theta}_{n} and theorem 17.1, it is clear why the limit is χ12\chi^{2}_{1}.

  • •
    ​

    In use: compute WW from the two maximised log-likelihoods, count r=dimΘ−dimΘ0r=\dim\Theta-\dim\Theta_{0}, reject if W≥χr,1−α2W\geq\chi^{2}_{r,1-\alpha}. The approximation can be poor for small nn, the test fails if the regularity conditions are violated, and no information about power is available.

Exercise 23.1.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. N⁢(μ,σ2)N(\mu,\sigma^{2}), where both parameters are unknown. Consider the hypothesis H0:μ=μ0H_{0}\colon\mu=\mu_{0} against the alternative H1:μ≠μ0H_{1}\colon\mu\neq\mu_{0}, where σ2\sigma^{2} is a nuisance parameter. Let Q=∑i(Xi−X¯)2Q=\sum_{i}(X_{i}-\bar{X})^{2}, S2=Q/(n−1)S^{2}=Q/(n-1) and t=n⁢(X¯−μ0)/St=\sqrt{n}\,(\bar{X}-\mu_{0})/S.

  1. 1.
    ​

    Following example 5.4, show that the maximised log-likelihood function over Θ\Theta is given by

    −n2⁢log⁡(2⁢π⁢Q/n)−n2.-\frac{n}{2}\log(2\pi Q/n)-\frac{n}{2}.
  2. 2.
    ​

    Show that, restricting ourselves to the region μ=μ0\mu=\mu_{0}, the maximum of the likelihood function is attained for σ^02=1n⁢∑i(Xi−μ0)2\hat{\sigma}_{0}^{2}=\frac{1}{n}\sum_{i}(X_{i}-\mu_{0})^{2} and find expressions for σ^02\hat{\sigma}_{0}^{2} in terms of QQ and X¯−μ0\bar{X}-\mu_{0} using the identity (2.1).

  3. 3.
    ​

    Deduce that

    W=n⁢log⁡(1+t2n−1),W=n\log\Bigl{(}1+\frac{t^{2}}{n-1}\Bigr{)},

    and that the generalised likelihood-ratio test rejects for large values of |t||t|.

  4. 4.
    ​

    State the approximate distribution of WW under H0H_{0} given by theorem 23.3 and justify your answer for the degrees of freedom. Recall that, under H0H_{0}, the statistic tt follows the tn−1t_{n-1} distribution by definition A.4 and corollary 10.10. For n=10n=10, determine the value of |t||t| above which the approximate test with critical value χ1,0.952=3.841\chi^{2}_{1,0.95}=3.841 rejects and compare your result with the 0.9750.975-quantile 2.2622.262 of the t9t_{9} distribution. Comment upon your result.

Exercise 23.2.

An experiment with mm possible outcomes is repeated nn times independently. The number of times each outcome occurs is given by NjN_{j} for outcome jj. Thus, (N1,…,Nm)(N_{1},\dots,N_{m}) is multinomially distributed with probability weights

ℙp⁢(N1=n1,…,Nm=nm)=n!n1!⁢⋯⁢nm!⁢p1n1⁢⋯⁢pmnm,n1+⋯+nm=n,\mathbb{P}_{p}(N_{1}=n_{1},\dots,N_{m}=n_{m})=\frac{n!}{n_{1}!\cdots n_{m}!}\,% p_{1}^{n_{1}}\cdots p_{m}^{n_{m}},\qquad n_{1}+\dots+n_{m}=n,

where p=(p1,…,pm)p=(p_{1},\dots,p_{m}) satisfies pj>0p_{j}>0 and ∑jpj=1\sum_{j}p_{j}=1. We test H0:p=p0H_{0}\colon p=p^{0} for a given vector p0=(p10,…,pm0)p^{0}=(p_{1}^{0},\dots,p_{m}^{0}) against the alternative p≠p0p\neq p^{0}.

  1. 1.
    ​

    Show that the log-likelihood function is ∑jnj⁢log⁡pj\sum_{j}n_{j}\log p_{j} up to an additive constant and use the Lagrange multiplier method, with constraint ∑jpj=1\sum_{j}p_{j}=1, to show that the maximum of the likelihood function is at p^j=nj/n\hat{p}_{j}=n_{j}/n.

  2. 2.
    ​

    Show that

    W=2⁢∑j=1mnj⁢log⁡njn⁢pj0,W=2\sum_{j=1}^{m}n_{j}\log\frac{n_{j}}{n\,p_{j}^{0}},

    where 0⁢log⁡0=00\log 0=0.

  3. 3.
    ​

    Why does theorem 23.3 state that W→dχm−12W\xrightarrow{\ \mathrm{d}\ }\chi^{2}_{m-1} under H0H_{0}?

  4. 4.
    ​

    A die is thrown 6060 times and the six faces appear 1313, 88, 1111, 66, 1212 and 1010 times, respectively. Test the hypothesis that the die is fair at level 0.050.05, using χ5,0.952=11.07\chi^{2}_{5,0.95}=11.07.

  5. 5.
    ​

    Let nj=n⁢pj0+djn_{j}=np_{j}^{0}+d_{j} so that ∑jdj=0\sum_{j}d_{j}=0. Show that, by expanding the logarithm to second order in dj/(n⁢pj0)d_{j}/(np_{j}^{0}), we find

    W≈∑j=1m(nj−n⁢pj0)2n⁢pj0.W\approx\sum_{j=1}^{m}\frac{(n_{j}-np_{j}^{0})^{2}}{np_{j}^{0}}.

    This is Pearson’s chi-squared statistic. Compute the value for the die in the previous part of the question.

Exercise 23.3.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. with Bernoulli distribution, i.e. with success probability p∈(0,1)p\in(0,1). Furthermore, let S=∑iXiS=\sum_{i}X_{i} and p^=S/n\hat{p}=S/n and consider the hypothesis H0:p=p0H_{0}\colon p=p_{0} against the alternative H1:p≠p0H_{1}\colon p\neq p_{0}.

  1. 1.
    ​

    Show that

    W=2⁢(S⁢log⁡p^p0+(n−S)⁢log⁡1−p^1−p0),W=2\Bigl{(}S\log\frac{\hat{p}}{p_{0}}+(n-S)\log\frac{1-\hat{p}}{1-p_{0}}\Bigr{% )},

    where we use the convention 0⁢log⁡0=00\log 0=0.

  2. 2.
    ​

    Approximately determine the distribution of WW under H0H_{0} for large nn and use the resulting distribution to derive a test at level α\alpha.

  3. 3.
    ​

    In n=50n=50 trials, we observed 3232 successes. Test the hypothesis H0:p=1/2H_{0}\colon p=1/2 at level 0.050.05 and determine the approximate pp-value.

  4. 4.
    ​

    In exercise 17.1, the model was reparametrised using the log-odds ψ=log⁡(p/(1−p))\psi=\log\bigl{(}p/(1-p)\bigr{)}. Show that the generalised likelihood-ratio statistic for testing H0:ψ=ψ0H_{0}\colon\psi=\psi_{0} coincides with the one in part (a) for p0=eψ0/(1+eψ0)p_{0}=e^{\psi_{0}}/(1+e^{\psi_{0}}), the value of the null proportion.