Lecture 20 Hypothesis Testing

So far we have asked questions about the value of θ\theta. In this lecture we will consider the second question from lecture 1: if there is a claim about θ\theta, is it supported by the data? We will spend some time to carefully set up the language of hypothesis testing, which will be used in lectures 22, 23, 25 and 29. We will conclude the lecture by considering two examples, both about the mean of a normal sample.

20.1 Hypotheses

Throughout, X1,…,XnX_{1},\dots,X_{n} follow a parametric model in the sense of definition 1.1. Questions like the one about fairness of a coin or about the effect of a treatment ask whether the true parameter value θ0\theta_{0} belongs to a subset of the parameter space Θ\Theta.

Definition 20.1.

A hypothesis is a statement of the form θ∈Θ′\theta\in\Theta^{\prime} for a subset Θ′⊆Θ\Theta^{\prime}\subseteq\Theta. In a testing problem, the null hypothesis H0:θ∈Θ0H_{0}\colon\theta\in\Theta_{0} and the alternative hypothesis H1:θ∈Θ1H_{1}\colon\theta\in\Theta_{1} are given, where Θ0\Theta_{0} and Θ1\Theta_{1} are disjoint subsets of Θ\Theta. A hypothesis is called simple, if the corresponding subset consists of a single point, and composite otherwise.

As an abbreviation, from this lecture on we write θ0\theta_{0} for the value chosen to test a null hypothesis. In lectures 16 and 17, the same symbol was used to denote the unknown, true value of the parameter. The context should make it clear which meaning is intended.

In most problems we have Θ1=Θ∖Θ0\Theta_{1}=\Theta\setminus\Theta_{0}, i.e. we can assume that exactly one of the two hypotheses is true. For example, for the mean θ\theta of a normal sample, the hypothesis H0:θ=θ0H_{0}\colon\theta=\theta_{0} is simple, whereas the one-sided alternative H1:θ>θ0H_{1}\colon\theta>\theta_{0}, the two-sided alternative H1:θ≠θ0H_{1}\colon\theta\neq\theta_{0} and the one-sided null hypothesis H0:θ≤θ0H_{0}\colon\theta\leq\theta_{0} are all composite.

The two hypotheses have different roles: the null hypothesis is the default position, and the alternative is the claim which requires evidence before we can act on it. Section 20.3 will make this distinction precise.

20.2 Tests, Critical Regions and Errors

A test decides between the two hypotheses under consideration, and can be described by the set of data values where it rejects H0H_{0}.

Definition 20.2.

A test of H0H_{0} against H1H_{1} is a rule which for every possible data vector x=(x1,…,xn)x=(x_{1},\dots,x_{n}) decides whether to “reject H0H_{0}” or “do not reject H0H_{0}”. The set

C={x|the test rejects H0 for the data x}C=\bigl{\{}\,x\mathrel{\big{|}}\text{the test rejects $H_{0}$ for the data $x$% }\,\bigr{\}}

is called the critical region (or rejection region) of the test. Alternatively, a test can be described by its test function φ⁢(x)=1\varphi(x)=1 for x∈Cx\in C and φ⁢(x)=0\varphi(x)=0 otherwise, where 11 indicates rejection.

In practice, the critical region is almost always of the form C={x∣T⁢(x)>c}C=\{x\mid T(x)>c\} or C={x∣T⁢(x)≥c}C=\{x\mid T(x)\geq c\} for a statistic TT given by definition 1.3, called the test statistic, and a constant cc called the critical value. The two forms of the critical region differ if T⁢(x)=cT(x)=c has positive probability (exercise 20.2 shows an example of this case, for discrete distributions).

A test applied to random data will sometimes wrongly decide. This can happen in two different ways.

Definition 20.3.

A type I error occurs, if the test wrongly rejects H0H_{0} when H0H_{0} is true. A type II error occurs, if the test wrongly does not reject H0H_{0} when H1H_{1} is true.

Table 20.1 summarises the possible outcomes of statistical tests, in relation to the truth of the hypothesis being tested.

do not reject H0H_{0} reject H0H_{0}
H0H_{0} true correct decision type I error
H1H_{1} true type II error correct decision (power)
Table 20.1: The four possible outcomes of a test.

To understand the interpretation of this table, we can use the following example: in a criminal trial, the null hypothesis is the presumption of innocence of the defendant, and the alternative is guilt. If an innocent person is convicted, this is a type I error. If a guilty person is acquitted, this is a type II error. The standard of proof “beyond reasonable doubt” is such that many acquittals of the guilty are acceptable in order to avoid too many wrongful convictions. The reason for this is that the two types of error have different consequences: the consequences of type I errors are more serious. Furthermore, since an acquittal does not prove innocence, we write “do not reject H0H_{0}” instead of “accept H0H_{0}” in the table.

20.3 The Power Function, Size and Level

The probabilities of the two errors depend on the true parameter value, and both can be found using a single function of θ\theta.

Definition 20.4.

The power function of a test with critical region CC is given by

β⁢(θ)=ℙθ⁢(X∈C)=𝔼θ⁢(φ⁢(X)),θ∈Θ,\beta(\theta)=\mathbb{P}_{\theta}(X\in C)=\mathbb{E}_{\theta}\bigl{(}\varphi(X% )\bigr{)},\qquad\theta\in\Theta,

i.e. the probability of rejecting H0H_{0} when the true parameter value is θ\theta. For θ∈Θ1\theta\in\Theta_{1} the value β⁢(θ)\beta(\theta) is called the power of the test at the alternative θ\theta.

For θ∈Θ0\theta\in\Theta_{0} the value β⁢(θ)\beta(\theta) is the probability of a type I error, and for θ∈Θ1\theta\in\Theta_{1} the value 1−β⁢(θ)1-\beta(\theta) is the probability of a type II error. The two probabilities cannot both be reduced by making the critical region larger or smaller. If the size of the critical region is increased, the value β⁢(θ)\beta(\theta) increases for all θ\theta, i.e. type I errors become more likely and type II errors become less likely. If the size of the critical region is decreased, the opposite happens. For this reason, we have to choose one of the two types of error to control and it transpires that we can decide which type of error to control by considering the asymmetry between the two types of hypotheses: type I errors are errors where the default position is abandoned without good reason, and thus we should restrict the probability of type I errors first and only then should we consider the power of the test.

Definition 20.5.

The size of a test with power function β\beta is given by

supθ∈Θ0β⁢(θ),\sup_{\theta\in\Theta_{0}}\beta(\theta),

where the supremum is the maximum probability of type I error among all possible null hypotheses. For α∈(0,1)\alpha\in(0,1), the test is a level α\alpha test, if the size of the test is less than or equal to α\alpha.

For a simple null hypothesis H0:θ=θ0H_{0}\colon\theta=\theta_{0}, the size of the test equals β⁢(θ0)\beta(\theta_{0}). The number α\alpha is called the significance level and the values 0.050.05 and 0.010.01 are only conventions. Among all level α\alpha tests, we prefer the test with the largest power at the alternatives. In lecture 22 we will see how to find the test with optimal power, when both the null and the alternative hypothesis are simple.

20.4 Testing the Mean of a Normal Sample

We now construct a test from first principles for X1,…,XnX_{1},\dots,X_{n} i.i.d. N⁢(θ,σ2)N(\theta,\sigma^{2}), where the variance σ2\sigma^{2} is known. Here we write Φ\Phi for the standard normal distribution function and zqz_{q} for the qq-quantile of the standard normal distribution, i.e. we have Φ⁢(zq)=q\Phi(z_{q})=q. The most commonly used values of this function are z0.95=1.645z_{0.95}=1.645 and z0.975=1.96z_{0.975}=1.96.

Example 20.6.

We want to test H0:θ=θ0H_{0}\colon\theta=\theta_{0} against H1:θ>θ0H_{1}\colon\theta>\theta_{0} at significance level α\alpha. If X¯\bar{X} is significantly larger than θ0\theta_{0}, we can conclude that the alternative is more likely to be true. Thus, we can choose X¯\bar{X} as the test statistic and reject if X¯>c\bar{X}>c. From proposition A.3 we know that under H0H_{0} the standardised mean n⁢(X¯−θ0)/σ\sqrt{n}\,(\bar{X}-\theta_{0})/\sigma is standard normally distributed. Thus, the probability of type I errors is

β⁢(θ0)=ℙθ0⁢(X¯>c)=ℙθ0⁢(n⁢(X¯−θ0)σ>n⁢(c−θ0)σ)=1−Φ⁢(n⁢(c−θ0)σ).\beta(\theta_{0})=\mathbb{P}_{\theta_{0}}(\bar{X}>c)=\mathbb{P}_{\theta_{0}}% \Bigl{(}\frac{\sqrt{n}\,(\bar{X}-\theta_{0})}{\sigma}>\frac{\sqrt{n}\,(c-% \theta_{0})}{\sigma}\Bigr{)}=1-\Phi\Bigl{(}\frac{\sqrt{n}\,(c-\theta_{0})}{% \sigma}\Bigr{)}.

Setting this probability equal to α\alpha, we find the critical value as n⁢(c−θ0)/σ=z1−α\sqrt{n}\,(c-\theta_{0})/\sigma=z_{1-\alpha} and thus the level α\alpha test rejects H0H_{0} if

equation (20.1) (20.1)
X¯>θ0+z1−α⁢σn,equivalently ifZ=n⁢(X¯−θ0)σ>z1−α.\bar{X}>\theta_{0}+z_{1-\alpha}\,\frac{\sigma}{\sqrt{n}},\qquad\text{% equivalently if}\qquad Z=\frac{\sqrt{n}\,(\bar{X}-\theta_{0})}{\sigma}>z_{1-% \alpha}.

This is the one-sided zz-test. Since n⁢(X¯−θ)/σ\sqrt{n}\,(\bar{X}-\theta)/\sigma is standard normally distributed under ℙθ\mathbb{P}_{\theta}, we can use the same standardisation to determine the power of the test. The power function is

equation (20.2) (20.2)
β⁢(θ)=ℙθ⁢(n⁢(X¯−θ)σ>z1−α−n⁢(θ−θ0)σ)=Φ⁢(n⁢(θ−θ0)σ−z1−α).\beta(\theta)=\mathbb{P}_{\theta}\Bigl{(}\frac{\sqrt{n}\,(\bar{X}-\theta)}{% \sigma}>z_{1-\alpha}-\frac{\sqrt{n}\,(\theta-\theta_{0})}{\sigma}\Bigr{)}=\Phi% \Bigl{(}\frac{\sqrt{n}\,(\theta-\theta_{0})}{\sigma}-z_{1-\alpha}\Bigr{)}.

Since the function Φ\Phi is increasing, the graph of β\beta is an S-shaped curve which starts at 0 and converges to 11, passing through the point (θ0,α)(\theta_{0},\alpha). Close to θ0\theta_{0} the power is only slightly larger than α\alpha, but as nn increases, the curve gets steeper. This fact is used in exercise 20.1 to determine the required sample size.

For n=25n=25, σ=2\sigma=2, θ0=10\theta_{0}=10 and α=0.05\alpha=0.05 the test rejects if X¯>10+1.645⋅2/5=10.658\bar{X}>10+1.645\cdot 2/5=10.658. At the alternative θ=11\theta=11 the power is Φ⁢(2.5−1.645)=Φ⁢(0.855)=0.804\Phi(2.5-1.645)=\Phi(0.855)=0.804, at θ=10.5\theta=10.5 it is only Φ⁢(1.25−1.645)=Φ⁢(−0.395)=0.346\Phi(1.25-1.645)=\Phi(-0.395)=0.346: a shift of half a standard deviation is detected in barely a third of all samples.

Since the power function (20.2) is increasing, the same test also has size supθ≤θ0β⁢(θ)=β⁢(θ0)=α\sup_{\theta\leq\theta_{0}}\beta(\theta)=\beta(\theta_{0})=\alpha for the composite null hypothesis H0:θ≤θ0H_{0}\colon\theta\leq\theta_{0}, attained at the boundary point; this is why definition 20.5 takes a supremum.

Example 20.7.

Assume that we want to test for both directions, i.e. we test H0:θ=θ0H_{0}\colon\theta=\theta_{0} against H1:θ≠θ0H_{1}\colon\theta\neq\theta_{0} and we reject for |X¯−θ0|>c|\bar{X}-\theta_{0}|>c. Under H0H_{0} the test statistic Z=n⁢(X¯−θ0)/σZ=\sqrt{n}\,(\bar{X}-\theta_{0})/\sigma is standard normally distributed and, since the normal distribution is symmetric, we have

ℙθ0⁢(|X¯−θ0|>c)=ℙθ0⁢(|Z|>n⁢cσ)=2⁢(1−Φ⁢(n⁢cσ)).\mathbb{P}_{\theta_{0}}\bigl{(}|\bar{X}-\theta_{0}|>c\bigr{)}=\mathbb{P}_{% \theta_{0}}\Bigl{(}|Z|>\frac{\sqrt{n}\,c}{\sigma}\Bigr{)}=2\Bigl{(}1-\Phi\Bigl% {(}\frac{\sqrt{n}\,c}{\sigma}\Bigr{)}\Bigr{)}.

Setting this equal to α\alpha, we find n⁢c/σ=z1−α/2\sqrt{n}\,c/\sigma=z_{1-\alpha/2} and thus the level α\alpha two-sided zz-test rejects H0H_{0} if

equation (20.3) (20.3)
|X¯−θ0|>z1−α/2⁢σn,equivalently if|Z|>z1−α/2.|\bar{X}-\theta_{0}|>z_{1-\alpha/2}\,\frac{\sigma}{\sqrt{n}},\qquad\text{% equivalently if}\qquad|Z|>z_{1-\alpha/2}.

Let δ=n⁢(θ−θ0)/σ\delta=\sqrt{n}\,(\theta-\theta_{0})/\sigma. Then, similar to the result in example 20.6, we find the power function

equation (20.4) (20.4)
β⁢(θ)=Φ⁢(−z1−α/2−δ)+1−Φ⁢(z1−α/2−δ).\beta(\theta)=\Phi(-z_{1-\alpha/2}-\delta)+1-\Phi(z_{1-\alpha/2}-\delta).

This function is symmetric about θ0\theta_{0} and has its minimum value α\alpha at this point. Thus, the graph of this power function is U-shaped instead of S-shaped. For n=25n=25, σ=2\sigma=2, θ0=10\theta_{0}=10 and α=0.05\alpha=0.05 the test rejects if |X¯−10|>1.96⋅2/5=0.784|\bar{X}-10|>1.96\cdot 2/5=0.784 and at θ=11\theta=11 we have δ=2.5\delta=2.5 and the power Φ⁢(−4.46)+1−Φ⁢(−0.54)=0.705\Phi(-4.46)+1-\Phi(-0.54)=0.705.

At θ=11\theta=11 the one-sided test has power 0.8040.804 against 0.7050.705, since it rejects all values with probability α\alpha from the upper tail. Thus, the one-sided test is superior to the two-sided test for all alternatives above θ0\theta_{0}, but is useless for alternatives below. Before we consider the data, we need to decide which departures from H0H_{0} are of interest. In lecture 22 we will see that there is no single test which is best for both directions, and in lecture 24 we will consider how size and power behave in simulation.

20.5 The pp-Value

A test at a fixed level only states whether H0H_{0} is rejected or not. The pp-value, introduced in this section, allows us to quantify the strength of the evidence in a more detailed way.

Definition 20.8.

Consider a family of tests with critical regions CαC_{\alpha} for all levels α∈(0,1)\alpha\in(0,1), such that Cα⊆Cα′C_{\alpha}\subseteq C_{\alpha^{\prime}} whenever α<α′\alpha<\alpha^{\prime}. Then the pp-value of the data xx is given by

p⁢(x)=inf{α∈(0,1)|x∈Cα},p(x)=\inf\bigl{\{}\,\alpha\in(0,1)\mathrel{\big{|}}x\in C_{\alpha}\,\bigr{\}},

the smallest level at which the data lead to rejection of H0H_{0}.

The tests from section 20.4 are nested in this way, since z1−αz_{1-\alpha} decreases as α\alpha increases. For a test which rejects for large values of a statistic TT and which has a simple null hypothesis, the smallest level at which the data can lead to rejection is

equation (20.5) (20.5)
p⁢(x)=ℙθ0⁢(T⁢(X)≥T⁢(x)).p(x)=\mathbb{P}_{\theta_{0}}\bigl{(}T(X)\geq T(x)\bigr{)}.

This is the probability under H0H_{0} of the test statistic taking a value at least as extreme as the observed value. For the one-sided zz-test this probability is p⁢(x)=1−Φ⁢(z)p(x)=1-\Phi(z), for the observed value z=n⁢(x¯−θ0)/σz=\sqrt{n}\,(\bar{x}-\theta_{0})/\sigma. For the two-sided test we get p⁢(x)=2⁢(1−Φ⁢(|z|))p(x)=2\bigl{(}1-\Phi(|z|)\bigr{)}. For the situation of example 20.6, the observed mean x¯=10.9\bar{x}=10.9 leads to z=2.25z=2.25 and thus to p=1−Φ⁢(2.25)=0.0122p=1-\Phi(2.25)=0.0122 for the one-sided test and p=0.0244p=0.0244 for the two-sided test. We can see that we can reject H0H_{0} at level α\alpha if and only if p⁢(x)≤αp(x)\leq\alpha, and thus the pp-value is a compact summary of a test. Under H0H_{0}, the pp-value of a continuous test statistic is uniformly distributed on (0,1)(0,1) (see exercise 20.4). This result is used in lecture 30.

Both statements assume that the test statistic is continuous. For discrete test statistics, e.g. the number of successes in a Bernoulli sample, the tail probability (20.5) can only take the countably many values ℙθ0⁢(T⁢(X)≥t)\mathbb{P}_{\theta_{0}}(T(X)\geq t) and the pp-value is discontinuous. Most levels α\alpha are not attained by any critical region. We can still perform the test for p⁢(x)≤αp(x)\leq\alpha and the resulting test is conservative. The size of the test is the largest attainable value which does not exceed the level α\alpha. The test will reject a true H0H_{0} less often than the nominal level suggests and the test will lose some power. The pp-value is no longer uniformly distributed either, but still satisfies ℙθ0⁢(p⁢(X)≤α)≤α\mathbb{P}_{\theta_{0}}(p(X)\leq\alpha)\leq\alpha. This is what allows the level to still be valid.

Since the pp-value is a probability, it is tempting to interpret the results in wrong ways; the following statements about a pp-value of 0.01220.0122 are all wrong.

  • •
    ​

    “The probability that H0H_{0} is true is 0.01220.0122.” Since θ\theta is a fixed, unknown number, H0H_{0} is either true or false; the probability in (20.5) refers to the data. To get a probability for a hypothesis, we have to introduce a prior distribution for θ\theta, i.e. we have to follow the Bayesian approach from lecture 28.

  • •
    ​

    “The probability that the observed data could have been produced by chance is 0.01220.0122.” The pp-value corresponds to extreme values of TT, not to the probability of the data.

  • •
    ​

    “If we had observed a pp-value of 0.40.4, we would have known that H0H_{0} is true.” Data which is compatible with H0H_{0} may be compatible with many different alternatives: the test from example 20.6 has power only 0.3460.346 at θ=10.5\theta=10.5.

  • •
    ​

    “The effect is important, because pp is small.” If nn is large enough, any difference between θ\theta and θ0\theta_{0} will result in a tiny pp-value. Reports should include both the estimate and the standard error.

None of these mistakes invalidates the pp-value as a tool, but they do show that the pp-value answers a very specific question, and that the same value should not be used to answer different questions.

Summary.
  • •
    ​

    A testing problem contrasts a null hypothesis H0:θ∈Θ0H_{0}\colon\theta\in\Theta_{0} against an alternative H1:θ∈Θ1H_{1}\colon\theta\in\Theta_{1}. The hypothesis is simple, if the corresponding subset of Θ\Theta is a single point, and composite otherwise.

  • •
    ​

    A test is given by its critical region CC, typically of the form T⁢(x)>cT(x)>c for a test statistic TT and critical value cc.

  • •
    ​

    Type I errors are cases where H0H_{0} is true but is wrongly rejected by the test, type II errors are cases where H0H_{0} is false but is not (correctly) rejected. The power function β⁢(θ)=ℙθ⁢(X∈C)\beta(\theta)=\mathbb{P}_{\theta}(X\in C) gives the probability of type I errors for Θ0\Theta_{0} and the power of the test for Θ1\Theta_{1}.

  • •
    ​

    The size of a test is supΘ0β\sup_{\Theta_{0}}\beta. A level α\alpha test has size at most α\alpha. The level is chosen first, and then the power of the test can be considered. For example, for the normal mean with known variance, the one-sided and two-sided zz-tests can be derived.

  • •
    ​

    The pp-value is the smallest level at which the data lead to rejection of the hypothesis. Note that this is not the probability that H0H_{0} is true, and a large pp-value does not establish H0H_{0}.

Exercise 20.1.

Consider the one-sided zz-test from example 20.6, at level α\alpha, and let θ1>θ0\theta_{1}>\theta_{0} be the alternative and 1−γ∈(α,1)1-\gamma\in(\alpha,1) be the target power.

  1. 1.
    ​

    Show that the test has power at least 1−γ1-\gamma at θ1\theta_{1} if and only if

    n≥((z1−α+z1−γ)⁢σθ1−θ0)2.n\geq\Bigl{(}\frac{(z_{1-\alpha}+z_{1-\gamma})\,\sigma}{\theta_{1}-\theta_{0}}% \Bigr{)}^{2}.
  2. 2.
    ​

    For σ=2\sigma=2, θ1−θ0=0.5\theta_{1}-\theta_{0}=0.5, α=0.05\alpha=0.05 and target power 0.90.9 (so that z0.9=1.282z_{0.9}=1.282), determine the smallest sample size which achieves the target power.

  3. 3.
    ​

    How does the required sample size change, when the difference θ1−θ0\theta_{1}-\theta_{0} is halved? What happens when σ\sigma is doubled?

Exercise 20.2.

Let X1,…,X10X_{1},\dots,X_{10} be i.i.d. Bernoulli with success probability pp and consider testing H0:p=1/2H_{0}\colon p=1/2 against H1:p>1/2H_{1}\colon p>1/2 using the test which rejects if T=∑iXi≥cT=\sum_{i}X_{i}\geq c for an integer cc.

  1. 1.
    ​

    Write down the size of the test as a function of cc, using the binomial distribution of TT under H0H_{0}.

  2. 2.
    ​

    Determine the smallest cc such that the test has level 0.050.05. Compute the size of the test for this case. Justify your answer by showing that no choice of cc can lead to a size of exactly 0.050.05.

  3. 3.
    ​

    Determine the power of the test from the previous part when p=0.7p=0.7, p=0.8p=0.8 and p=0.9p=0.9. Comment on the capabilities and limitations of the test with ten observations.

  4. 4.
    ​

    Nine successes are observed. Compute the pp-value and state the decision at level 0.050.05.

Exercise 20.3.

Let X1,…,Xn∼Uniform⁢(0,θ)X_{1},\dots,X_{n}\sim\text{Uniform}(0,\theta) be i.i.d. as in example 1.2 and let M=maxi⁡XiM=\max_{i}X_{i}, where the density was found in example 8.10. We test H0:θ=θ0H_{0}\colon\theta=\theta_{0} against H1:θ>θ0H_{1}\colon\theta>\theta_{0}.

  1. 1.
    ​

    Show that ℙθ⁢(M≤m)=(m/θ)n\mathbb{P}_{\theta}(M\leq m)=(m/\theta)^{n} for 0≤m≤θ0\leq m\leq\theta.

  2. 2.
    ​

    Consider the test which rejects if M>cM>c, for a constant c≤θ0c\leq\theta_{0}. Determine the size of this test and the value of cc which makes the size equal to α\alpha.

  3. 3.
    ​

    Determine the power function of the test from the previous part, and evaluate this function at n=5n=5, θ0=1\theta_{0}=1, α=0.05\alpha=0.05 and θ=1.2\theta=1.2.

  4. 4.
    ​

    Consider the second test, which rejects if M>θ0M>\theta_{0}. Show that this test has size 0. Determine the power function of this test and compare the power functions of the two tests. Which of the two tests would you choose and why?

Exercise 20.4.

A sample of size n=16n=16 from N⁢(θ,σ2)N(\theta,\sigma^{2}) with known σ=4\sigma=4 has mean x¯=52.3\bar{x}=52.3. We want to test θ0=50\theta_{0}=50.

  1. 1.
    ​

    Determine the pp-value of the one-sided zz-test with H0:θ=50H_{0}\colon\theta=50 against H1:θ>50H_{1}\colon\theta>50 and state the decisions for 0.050.05 and 0.010.01.

  2. 2.
    ​

    Determine the pp-value of the two-sided test against H1:θ≠50H_{1}\colon\theta\neq 50 and state the same two decisions.

  3. 3.
    ​

    Your colleague summarises the first part of the question as “there is a 1.1%1.1\% chance that the mean is 5050”. What is wrong with this sentence? Write one correct sentence to summarise the first part of the question.

  4. 4.
    ​

    Show that under H0H_{0} the pp-value of the one-sided zz-test, p⁢(X)=1−Φ⁢(Z)p(X)=1-\Phi(Z) with Z=n⁢(X¯−θ0)/σZ=\sqrt{n}\,(\bar{X}-\theta_{0})/\sigma, satisfies ℙθ0⁢(p⁢(X)≤u)=u\mathbb{P}_{\theta_{0}}\bigl{(}p(X)\leq u\bigr{)}=u for all u∈(0,1)u\in(0,1).