Appendix C Solutions to the Exercises
This appendix contains worked solutions to the exercises found at the end of each lecture. The solutions are grouped by lecture and refer to the exercises by number.
C.1 Statistical Models and the Estimation Problem
Here we solve the exercises at the end of lecture 1.
-
1.
Since the tosses are independent and since each toss has probability of being heads, we can describe as a sequence of independent and identically distributed Bernoulli trials with probability weights for and parameter space .
-
2.
We can use the proportion of heads in the sample, , as an estimator.
-
3.
The value is a function of the random variables before the coin tosses occur, so the value is still indeterminate. After the tosses have been made and the numbers are known, is a function of numbers, i.e. a single number.
∎
A statistic is a function of the data which does not use the unknown parameter .
-
1.
The range is computed from the data alone, so it is a statistic.
-
2.
The value contains the unknown mean and thus cannot be computed from the data, so it is not a statistic. If was known, then it would be a statistic.
-
3.
The sample variance uses the sample mean , which is computed from the data, in place of . Thus, it is a statistic.
∎
-
1.
Since the observations are independent, the joint density is the product of the marginal densities:
whenever all observations are in the support, i.e. when for all . Otherwise, the density is zero. This can also be written as the density being non-zero only when and .
-
2.
Each observation satisfies with probability one, since the support is . The maximum of numbers, none of which are larger than , cannot exceed . Thus, .
-
3.
Since the distribution of each is continuous, we have for all . The maximum equals only if at least one observation equals this value, and thus we get
Together with the first part of the solution this shows that with probability one. Thus, the random variable is non-negative and is zero with probability zero. A non-negative random variable with expectation zero equals zero with probability one, and thus we find , which is the same as . Thus, the estimator never overestimates and underestimates on average.
∎
C.2 Bias, Mean Squared Error and Consistency
Here we solve the exercises at the end of lecture 2.
-
1.
Since , we can use lemma 2.2 to conclude . Thus we have for all and the estimator is unbiased.
- 2.
∎
-
1.
Using the given density
we find the bias to be . Since this converges to zero as , the estimator is asymptotically unbiased. Since the bias is negative, the maximum systematically underestimates , as expected from exercise 1.3.
- 2.
-
3.
The mean squared error is of order , while the mean squared error of the sample mean is of order . Thus, the maximum converges to the target faster than the average does. The very fast convergence is caused by the support of the uniform model changing as a function of the parameter. This effect, mentioned in example 1.2, will be seen again in lecture 13.
∎
Since , we have and using the gamma density of we get
For , the remaining integral can be evaluated as the gamma integral . Thus we get
The bias is given by
Since this is strictly positive, we see that systematically overestimates on average and thus is biased. As , the bias converges to zero and thus is asymptotically unbiased. This completes the solution. ∎
C.3 Likelihood and the Method of Moments
Here we solve the exercises at the end of lecture 4.
-
1.
Using the Poisson weights from table 1.1, we find the likelihood to be
Taking logarithms we get
as claimed.
-
2.
Taking derivatives with respect to , and noting that the last term does not depend on , we find the score function to be
Setting this equal to zero and solving gives . Thus, the score function is zero at the sample mean.
-
3.
For the Poisson distribution we have . Thus, the only moment equation we have is : the method of moments estimator is the sample mean. This is the value at which the score function is zero, i.e. for the Poisson distribution the method of moments coincides with the maximum likelihood estimator from lecture 5.
∎
-
1.
Since we have and the parameter is scalar, we get the moment equation . Solving this equation we find
-
2.
Since each observation satisfies , we have , with equality if and only if all observations equal one. Thus, the interval of values where the estimate is possible is . Since for every the dataset with values in has positive probability, no estimate in this range is incompatible with the data (this is in contrast to the example of the uniform distribution in the text).
∎
-
1.
From table A.2 we know that and . Thus, the first population moment is . Rewriting the relation we find
as claimed.
-
2.
Using the moments from the first part, the moment equations from definition 4.6 are
From the first equation we find directly. Substituting this value into the second equation we find
where the last equality sign comes from expanding the square and using the relation . These are the estimators given in example 4.8 and thus the solution is complete.
∎
- 1.
-
2.
The maximum has smaller mean squared error, if and only if
This inequality is equivalent to , i.e. to for all . (The equality signs hold for and .) Thus, for three or more observations the biased maximum is better than the unbiased method of moments estimator, and the rate of the maximum makes the difference increase quickly with . This is another example of the lesson from lecture 2 that unbiasedness is not sufficient: the moment estimator squashes the data into the sample mean, whereas the maximum uses the fact that the support of the distribution moves with the parameter.
∎
-
1.
The distribution depends on only via . Thus, for different parameter values and , the distribution is the same, e.g. and both lead to . Thus, the model is not identifiable.
-
2.
To make the model identifiable, we can restrict ourselves to : for the distributions and have different variances and thus are different. (The other half line would work equally well.)
-
3.
Since the first moment for all , we cannot use this moment to derive any information about the parameter. We can use the second moment instead. Since , the moment equation is and taking the positive root, as required for , we get
This completes the solution.
∎
C.4 Maximum Likelihood Estimation
Here we solve the exercises at the end of lecture 5.
-
1.
Using the geometric weights, the likelihood function is
Taking logarithms we get , as claimed.
-
2.
Taking derivatives we find the score as
Setting this equal to zero and clearing the denominators we find , i.e. . The only solution to this equation is . Since all observations satisfy , we have and thus
for all . Thus, is strictly concave and the root is the global maximum. Thus, the MLE is . (If all observations equal one, i.e. if , then is strictly increasing and no maximiser exists in the open interval . The same boundary phenomenon as for the Bernoulli distribution in the text occurs. In this case, the formula gives the boundary value .)
-
3.
In exercise 4.2 we have found the method of moments estimator to be . Thus, for the geometric distribution, as for the Bernoulli, Poisson and exponential distributions, the two different construction methods coincide.
∎
-
1.
The log-likelihood function is
and the score function is
Setting the score equal to zero, we find and thus the only root is . Since the score is positive for and negative for , this root is the global maximum and the MLE is .
- 2.
- 3.
∎
-
1.
Since and since has gamma density on , we get the same integral as in exercise 2.3, but with one fewer power of , i.e. for we have
as claimed.
-
2.
Expanding the square and using the linearity of the expectation, we find
In the text we have seen that , and using this value and the second moment from the previous part of the question, we get
For the MLE we set and thus we get
For the unbiased estimator we set and thus we get the mean squared error as the variance
Dividing the two expressions we get , which is always smaller than for .
-
3.
The expression inside the brackets is a quadratic function of with positive leading coefficient, so the minimum is at the point where the derivative is zero. This point is
Substituting this value back into the expression inside the brackets we find
Since , this value is better than the unbiased estimator () and the MLE (), as shown in the previous part of the question. Since the winning estimator is smaller than the MLE, we know that the estimator is biased downwards. This is the same effect as we have seen for the normal variance in section 2.3, where the optimal divisor was instead of the unbiased value or the value found by maximum likelihood, i.e. : the mean squared error can be minimised by a deliberately biased estimator, and the amount of bias depends on the family of distributions under consideration. There is no one-size-fits-all solution for this kind of problem.
∎
-
1.
Up to the factor , which we can cancel, the score is . Multiplying through by the positive quantity , the likelihood equation becomes
as claimed.
-
2.
For we have and , and up to an additive constant the log-likelihood is . The product inside the logarithms is
which is the polynomial and thus .
-
3.
Since the logarithm is increasing, takes its largest value when is smallest. Let . Then we have , i.e. we get a quadratic function of with minimum at . If , this minimum is at a negative value of and thus is increasing on and takes its minimum at . The midpoint is then the only maximiser of and, since the factor from the first part of the solution has no real zero, is the only root of the likelihood equation. If , then takes its minimum at , i.e. at which are the zeros of . These two values are the maximum likelihood estimates, and since is even in , they have equal likelihood. Between these two values, at , the polynomial has a local maximum, because the coefficient of is negative. Thus, has a local minimum at the midpoint. This completes the proof.
∎
Since is a bijection, there exists an inverse function and we can re-index the model, using instead of , i.e. if we let the family of distributions is unchanged, but the likelihood function is . For any , the value falls into the set and by the definition of we have
Thus, is a maximum of the function on the set and consequently is a maximum likelihood estimator for . This completes the proof. ∎
-
1.
We have and thus the event is contained in the event . We get
i.e. the unconditional probability to survive beyond time . We can conclude from this that, if a lifetime has already passed time , the remaining lifetime has the same distribution as a fresh lifetime does. In other words, the time already elapsed does not change the distribution of the remaining lifetime.
-
2.
The conditional density of given the event is the density restricted to the set , normalised to be a density again. We get for . Thus we have
Using integration by parts, with and , we find and thus
Dividing by gives as required. This result is consistent with the first part of the solution: given survival until time , the remaining lifetime is distributed like a fresh exponential lifetime. Thus, the conditional expectation is the unconditional mean , plus the time already elapsed, giving . This completes the solution.
∎
-
1.
From the solution to exercise 4.1 we know that the score is given by and this score equals zero at . Let . Then the second derivative is for all and thus is strictly concave. This shows that the root is the unique global maximum and thus the MLE is given by .
-
2.
The map is continuous and strictly decreasing and therefore maps onto bijectively. Thus, by theorem 5.6, the MLE for is given by
-
3.
Using the independence of the observations and the given identity with we get
Since for all , we can choose to get and thus . This gives
and thus the estimator systematically overestimates on average and is biased for all . On the other hand, expanding the exponential we find as and thus and the estimator is asymptotically unbiased. There is no contradiction to the unbiasedness of : unbiasedness is not preserved under the nonlinear map , as discussed after theorem 5.6, whereas the maximum likelihood property is.
∎
-
1.
Solving the equation for we find and taking logarithms we find . The map is continuous and strictly decreasing on and since as and as , it maps onto itself in a bijective way.
- 2.
-
3.
From lemma 2.2 we know and thus for all and thus is unbiased. There is no contradiction: unbiasedness is a property of a parametrisation and is in general lost under nonlinear transformations of the parameter. Thus, the bias of the MLE for the rate is not relevant for the bias of the estimator for the median, which is a linear function of .
∎
C.5 Exponential Families
- 1.
-
2.
For the geometric series gives
but for the series diverges. Thus we have and
Substituting , so that , we get as required. As ranges over , the natural parameter ranges over the entire open interval , and thus the family has full rank.
- 3.
∎
- 1.
-
2.
Using the binomial theorem, we find
for all . Thus, and . This is times the cumulant function of the Bernoulli distribution.
- 3.
∎
-
1.
We can write and, collecting all factors into one exponential, for we find
This is the form of definition 7.8 with , with for and otherwise, natural statistic , natural parameter and .
-
2.
As varies over , the first coordinate varies over and as varies over , the second coordinate varies over , independently of each other. Thus, the set of natural parameter values is the open quadrant . This is a non-empty open subset of and the family has full rank as in definition 7.10.
-
3.
Knowing , we can include the factor into the function : for we have
This is a one-parameter exponential family in the sense of definition 7.1, with for , natural statistic and natural parameter .
∎
- 1.
-
2.
For and we have to evaluate
We can complete the square to find and thus
since the remaining integrand is the density of and thus integrates to one. The integral is finite for all and thus and . Taking derivatives we find and and substituting in theorem 7.7 we get and , as expected. Finally, , which is the function from example 7.9.
∎
- 1.
-
2.
Differentiating once more and using theorem 7.7 again, we find
for all . Thus is strictly concave on the interval , and a strictly concave function has at most one critical point, which is then its global maximum. Any root of the likelihood equation is therefore the unique maximiser of , i.e. the unique MLE for .
-
3.
For the Poisson distribution we have and , and thus the likelihood equation reads , with root whenever . Since is a bijection between and , theorem 5.6 gives , the MLE of exercise 5.7. For the Bernoulli distribution we have and , and thus the likelihood equation reads , with root whenever . Since is a bijection between and , theorem 5.6 gives , the MLE of example 5.3. The restriction is the condition discussed after that example: when it fails the likelihood equation has no root and the MLE does not exist in the open parameter space.
-
4.
The equation of the first part says that the MLE is the parameter value under which the population mean of the natural statistic agrees with its sample mean, i.e. the MLE is the method of moments estimator based on the moment rather than on . Whenever , as for the Bernoulli, Poisson and exponential distributions and for the normal mean with known variance, the two moments are the same, and the MLE coincides with the method of moments estimator from definition 4.6; this is the coincidence observed in the remark at the end of the method of moments section of lecture 4. For the gamma distribution of exercise 7.3 the natural statistic is , and thus maximum likelihood matches the means of and , whereas the method of moments in example 4.9 matched the means of and . Different moments are matched, and thus the two estimators differ.
∎
-
1.
Writing the expectation as an integral (or a sum, for a discrete model), we find
where the last step uses the definition of at the point . Taking logarithms gives , as claimed.
-
2.
By the first part we have . Differentiating both sides with respect to , and taking the derivative inside the expectation on the right, we get
At we have , and thus and . On the other hand, from we get directly and . Comparing the two computations gives and , which is theorem 7.7.
- 3.
∎
- 1.
-
2.
Let and . As varies over the interval , the first coordinate also varies over the interval . Substituting for the second coordinate, we find
Thus, the set of natural parameter values is the curve . Any open subset of which contains a point on the curve also contains a small disc around this point, and thus points off the curve, obtained by slightly changing while keeping fixed. Thus the curve does not contain any non-empty open subsets of . By definition 7.10, the family is not of full rank, as claimed.
∎
C.6 Sufficiency
Here we solve the exercises at the end of lecture 8.
-
1.
If , the event is contained in the event and we have
-
2.
If , the events and are disjoint and the conditional probability is zero. Thus the result does not depend on and is sufficient for by definition 8.1.
- 3.
∎
-
1.
The density of a single observation is . Since the joint density is the product of these densities, and since a product of indicators can be written as the indicator of the intersection of the corresponding events, we find that all conditions are satisfied if and only if the smallest observation is greater than or equal to and the largest observation is less than or equal to . Thus we have
-
2.
Let and . Then depends on the data only through , the function does not depend on and from the first part of the solution we know that for all and . By theorem 8.2 the statistic is sufficient for .
-
3.
The factor depends on the data but does not depend on and thus can be included in . The factor depends on and thus a function involving this factor cannot be in . This factor must be included in and, since it only depends on the data through , the maximum is the sufficient statistic resulting from this factorisation. It is a common mistake to write the joint density as and to conclude that no additional statistic is needed, since the indicator is where the data and the parameter interact.
∎
- 1.
-
2.
For two data vectors and with non-negative entries, the first part gives
If , the ratio equals for every . If the sums differ by , the ratio is , which is a strictly monotone function of and thus not constant on . The ratio is therefore free of if and only if and by theorem 8.6 we find that is minimal sufficient.
-
3.
Since , the MLE is given by , a function of . For fixed data, the likelihood function is , where is a positive constant not depending on . Thus, the maximiser of over is the same as the maximiser of , which depends on the data only through . Whatever the MLE is, it can be written as a function of the sufficient statistic by the factorisation in theorem 8.2.
∎
- 1.
- 2.
-
3.
Using the expanded form of the joint density from the previous part, and writing and , the ratio of the densities of two data vectors is
since the factors and cancel. If , which is the statement , the ratio equals for all parameter values. Conversely, suppose the ratio does not depend on , so that the exponent equals a constant for all and all . Taking , the expression is constant in , which forces and then . With the exponent is , and this is constant in only if . Thus the ratio is free of the parameters if and only if , and is minimal sufficient by theorem 8.6.
-
4.
No. By the previous part is minimal sufficient, and thus, by definition 8.5, would have to be a function of if were sufficient. It is not: for the data sets and have the same sample mean , but and , so that . Two data sets with the same value of can therefore have non-proportional likelihood functions, and alone loses information about .
∎
-
1.
The exponential distribution with rate has distribution function for , and thus . Proposition 8.9 gives
which is the exponential density with rate . Directly, the minimum exceeds if and only if every observation does, and by independence
which is the probability that an exponential random variable with rate exceeds . The two computations agree.
-
2.
From table A.2, an exponential random variable with rate has mean and variance . Thus , so is unbiased for , and .
-
3.
By lemma 2.2 the sample mean has , and thus we find that equals , so that the estimator based on the minimum has times the variance of the sample mean. Worse, its variance does not decrease at all as grows. Indeed, by the first part has the exponential distribution with rate for every , and thus its sampling distribution does not change with the sample size, and is not a consistent estimator for . It is not a function of either: for the data sets and have the same sum but minima and . The estimator discards the information which the sufficient statistic retains, and lecture 10 shows how conditioning on a sufficient statistic repairs such an estimator.
∎
-
1.
Since and are both Bernoulli with success probability , we have for all and thus is an unbiased estimator for zero. The variable equals on the event and, since the two events are independent, this happens with probability . Thus we have .
-
2.
The function satisfies for all , but does not equal zero with probability one, and thus the condition from definition 8.8 is violated: the whole sample is not complete. Nevertheless, the statistic is sufficient, because once the whole sample is given, there is no variability left. This completes the proof.
∎
C.7 Rao–Blackwell and Lehmann–Scheffe
-
1.
Since takes only the values and , we have for all and thus is an unbiased estimator for . The indicator function is a Bernoulli random variable with success probability , and thus we have .
-
2.
Let where . By the given fact, is Poisson with parameter and is Poisson with parameter , and is independent of . For , the event is the event and thus we get
As expected, the parameter cancels since is sufficient. Given , the first observation can be seen as the number of successes in independent trials with success probability , and in the model each of the counted events is equally likely to have been any of the observations.
-
3.
The conditional expectation of an indicator function is a conditional probability. Thus, using the second part of the question, we find
-
4.
Using the probability generating function with , we find
This agrees with the second statement of the Rao–Blackwell theorem, theorem 10.1.
-
5.
From example 10.6 we know that the sum is a complete and sufficient statistic, and is a function of which is unbiased for . Thus, by the Lehmann–Scheffe theorem, theorem 10.4, is the UMVUE for . The MLE is also a function of , but in exercise 5.7 we have seen that , and thus it is biased and cannot be the UMVUE. Nevertheless, for large , the two estimators are close since and and thus : the UMVUE is obtained by applying a small downward correction to the MLE in order to compensate for the upward bias.
∎
- 1.
-
2.
The random variables and are independent, has density and has the quoted gamma density with , and thus the pair has joint density for . The map is linear with determinant one, and thus by the change of variable formula, proposition A.19, the pair has joint density
for , and zero otherwise. Dividing by the density of , we obtain the conditional density
for . Both and have cancelled, and thus the conditional density does not depend on , which is the sufficiency of seen from definition 8.1. For the density is on : given the total of two observations, the first is uniformly distributed over the possible values.
-
3.
Substituting , we find
Integrating by parts, with differentiated and integrated, gives
and thus . Consequently : Rao–Blackwellising the single observation , which has variance , gives the sample mean, which has variance , and the reduction promised by theorem 10.1 is by a factor of .
-
4.
Since the are i.i.d. and is a symmetric function of them, the pair has the same joint distribution for every , and thus the conditional expectations are the same random variable for all . Summing over and using the linearity of conditional expectation, we find
and thus . This argument uses only the symmetry of an i.i.d. sample and none of the exponential form, and so it applies to every model in which is sufficient.
∎
- 1.
- 2.
-
3.
From the solution to exercise 2.2 we know and thus
The variance of a single observation is , as shown in table A.2, and thus, using lemma 2.2, we find
The ratio , which is always less than or equal to one for all , equals one only for , where both estimators equal . For large , the variance of the UMVUE is of order , compared to the order for , similar to the result in exercise 2.2.
-
4.
From theorem 10.1, the Rao–Blackwell theorem, we know that the conditional expectation is a random variable, which is a function of and is unbiased for . Since is complete, the first step of the proof of theorem 10.4 shows that there can only be one such function and thus we can conclude with probability one. For we have and , which coincides with the formula .
∎
-
1.
Using the definition of the exponential family, we can write the density as
which is of the form given in definition 7.8 with , natural statistic and natural parameter
-
2.
As varies over , the second component can take any value in . For fixed , the first component can take any value in as varies. Thus, the range of is the open half-plane , which is an open, non-empty subset of and thus the family has full rank in the sense of definition 7.10.
-
3.
From corollary 8.3 and the first statement of proposition 10.5 we know that the natural statistic of the sample is complete and sufficient. From the pair we can compute and , and conversely and , and thus the map between the two pairs is one-to-one. Since a one-to-one function of a complete sufficient statistic is itself complete and sufficient, the pair is complete and sufficient.
∎
-
1.
Let , so that are i.i.d. whatever the value of . Since subtracting the same constant from every observation shifts the maximum and the minimum by that constant, we have
and the right-hand side is a function of random variables whose joint distribution does not involve . Thus the distribution of is the same under every , and is ancillary for in the sense of definition 10.8.
-
2.
In the proof of corollary 10.10 we have seen that, for known , the normal distribution is a full-rank exponential family with natural statistic , and thus is complete and sufficient for by corollary 8.3 and proposition 10.5. Since is ancillary by the first part, Basu’s theorem, theorem 10.9, shows that and are independent.
-
3.
Now let be known and unknown, and write with i.i.d. standard normal. Then , and thus for example is proportional to , and this expectation is strictly positive rather than zero, since and are independent and continuous and thus differ with probability one. The distribution of changes with the parameter, and is not ancillary for . The scaled range does have a distribution free of , but it is not a statistic in the sense of definition 1.3, because computing it requires the unknown . Quantities of this kind, whose distribution is known although they involve the parameter, return in lecture 25, where they are used to build confidence intervals.
∎
C.8 Fisher Information
-
1.
The score for the observation is given by . This takes values if (with probability ) and if (with probability ). Thus we have
as given in lemma 11.1.
- 2.
- 3.
∎
- 1.
-
2.
Taking derivatives again, we find and thus
On the other hand, since the score is plus a constant, we can find the variance as
Both approaches give the same result.
-
3.
Since the density can be written as
we have , , and . Thus, the model is an exponential family as in definition 7.1. The natural parameter takes values and the cumulant function is
where is the mean and is the variance, as predicted by theorem 7.7. Since , we find by proposition 11.9 . This is the same value as we obtained above.
-
4.
From proposition 11.4 we know that the information of the sample is and thus . Since is unbiased, its mean squared error equals its variance and thus the variance of the MLE is the reciprocal of the Fisher information for all sample sizes, as it was for the sample mean in exercise 11.1 and example 11.7. From lecture 13 we know that this is the smallest variance an unbiased estimator can have.
∎
-
1.
The log-likelihood of a single observation is and thus the score is given by
Using we find
-
2.
Taking derivatives again, we find and taking expectations we get
From proposition 11.4 we know that the sample has information .
-
3.
Since the score is plus a constant, we find
This agrees with the result from the previous part.
∎
-
1.
Since we have and using proposition 11.9 we find . This is the variance of the natural statistic , as expected.
- 2.
-
3.
We have with and using proposition 11.10 we find
or using the notation we find . As , the information goes to infinity. This makes sense, since converges to zero, and the information measures accuracy on the absolute scale of the parameter: a probability which is close to zero is estimated with small absolute error by any reasonable estimator, and the approximate variance from lecture 17 is tiny, even if the error relative to the size of is not small.
∎
- 1.
-
2.
Using proposition 11.4, in the -parametrisation, the information of the sample about is given by . From table A.2 we know that and thus we have , by lemma 2.2. The variance of the unbiased estimator for equals the reciprocal of the information of the sample, and from lecture 13 we know that no unbiased estimator can be better.
∎
The right-hand side of (11.1) with is an integral over the whole real line, but and thus for and for ; since is a single point, it does not contribute to the integral; only , where , is left. We have
while the left-hand side can be found as
The difference between the two sides is , which is the value of the density at the upper limit of the integral. The derivative of the integral picks up this boundary term (from the moving upper limit) in addition to the derivative of the integrand, but the differentiation under the integral sign ignores this boundary term. ∎
C.9 The Cramer–Rao Inequality
-
1.
For we have
and taking derivatives we find the density for . This is the exponential density with rate , i.e. the density from table A.2 and is by the statement quoted in the question. Thus we have .
- 2.
- 3.
- 4.
∎
- 1.
-
2.
The sample mean is an unbiased estimator for and using lemma 2.2 we find , since the variance of a Bernoulli distribution is . This bound coincides with the bound from the definition, and thus is efficient by definition 13.4, since the exchange identity (13.1) holds because the Bernoulli distribution is a regular exponential family and has finite variance, as noted after lemma 13.1.
- 3.
-
4.
For we have and thus using theorem 13.3 we get the bound
The Bernoulli distribution is an exponential family with natural statistic and . From the discussion after proposition 13.5 we know that only affine functions of the form can be used to construct an efficient estimator. Since the odds are not an affine function of , there can be no efficient estimator for the odds.
-
5.
Any estimator can take at most different values, one for each possible value of the sample , and the expectation of the estimator is
which is a polynomial in of degree at most . A polynomial function is bounded on , whereas as and thus no choice of makes the estimator equal to the odds for all . Thus there is no unbiased estimator for the odds, efficient or otherwise.
∎
-
1.
From example 11.6 we know that . Thus, the Cramer–Rao bound for unbiased estimators for is . Since the Poisson distribution has mean and variance , lemma 2.2 shows that and . Thus, is an unbiased estimator and satisfies the bound: it is an efficient estimator for , since the exchange identity (13.1) holds because the Poisson distribution is a regular exponential family and has finite variance, as noted after lemma 13.1.
-
2.
The log-likelihood function is and thus the score function is
Re-arranging this we find , which is condition (13.2) with , and ; indeed .
-
3.
From the discussion after proposition 13.5 we know that in a one-parameter exponential family, only affine functions of the mean of the natural statistic allow for efficient estimators. In this case we have and is not an affine function of . Thus, no efficient estimator for exists. There is no contradiction in this result: example 10.6 shows that has the smallest variance amongst all unbiased estimators for , but this smallest variance is strictly larger than the Cramer–Rao bound. Efficiency is a stronger property than being UMVUE, and in this case the bound cannot be achieved.
∎
-
1.
Example 11.8 gives , and thus and the bound is .
-
2.
Since is unbiased, its mean squared error is its variance, . The efficiency is
which is less than one for every and tends to one as : the estimator falls short of the bound for every finite sample size, but only by a factor which disappears in the limit.
-
3.
The exponential distribution with rate is an exponential family with , , and , as in example 7.4, and thus . By the discussion after proposition 13.5, an efficient estimator exists only for affine functions of . The rate is not of this form, and thus no unbiased estimator for is efficient. The same conclusion can be reached directly: condition (13.2) with would read , and since does not depend on , the coefficients and would have to be constants and , giving , which cannot equal for all .
-
4.
The natural parameter ranges over the open interval , and thus the family is of full rank and is complete and sufficient by corollary 8.3 and proposition 10.5. Since is a function of and unbiased for , the Lehmann–Scheffe theorem, theorem 10.4, shows that it is the UMVUE for . Together with the previous part this is an example of a UMVUE which is not efficient: the bound is not attained by any unbiased estimator, and the best unbiased estimator has variance .
-
5.
For we have , and theorem 13.3 gives the bound
The sample mean is unbiased for and has variance by lemma 2.2, and thus is efficient for , since the exchange identity (13.1) holds because the exponential distribution is a regular exponential family and has finite variance, as noted after lemma 13.1. This is consistent with the previous parts: is the mean of the natural statistic, for which proposition 13.5 guarantees an efficient estimator, whereas the rate is a non-affine function of it. Whether an efficient estimator exists depends on which function of the parameter we set out to estimate.
∎
-
1.
Since is given and is the parameter, the log-likelihood function is and thus
Since , we can use theorem 11.3 to find
Thus, the Cramer–Rao bound is .
-
2.
The are i.i.d. standard normally distributed, i.e. by definition A.2, with mean and variance . Thus, has mean and variance . This is the bound and thus is efficient, since the exchange identity (13.1) holds because the normal distribution with known mean is a regular exponential family and has finite variance, as noted after lemma 13.1. Alternatively, the score for the first part can be written as , which satisfies condition (13.2) with .
-
3.
From proposition A.3 we know that , with variance . Thus we have
Thus, the sample variance does not achieve the bound, and the efficiency of the sample variance is . Since the bound still holds for unknown , and since is the UMVUE for by example 10.7, this is a second example of a best unbiased estimator which is not efficient: no unbiased estimator for can achieve the bound if also needs to be estimated, and the loss of one degree of freedom is the price to be paid for not knowing .
∎
- 1.
-
2.
Since , the estimator is an unbiased estimator for the parameter , and theorem 13.2, applied in the -parametrisation, bounds its variance by . Using the first part we get
which is the bound of theorem 13.3. The two theorems thus give the same bound, and the bound does not depend on how the model is parametrised. This completes the proof.
∎
C.10 Multiparameter Fisher Information
-
1.
The Jacobian matrix of is given by
This matrix is invertible for . Since all three matrices are diagonal, the product is also diagonal and consists of the entries and . Thus, by proposition 14.4, we have .
-
2.
The log-likelihood function in terms of is
Taking derivatives we get and . Taking derivatives again, we find
Using and we find that the mixed derivative has expectation zero, and the last term has expectation . Thus, by proposition 14.3, the information matrix is given by . This agrees with the result from part (a).
-
3.
The inverse of the matrix from part (a) is given by . Thus, by theorem 14.5, any unbiased estimator for has variance of at least . The sample variance is an unbiased estimator for by lemma 2.3, but the square root of the sample variance is not an unbiased estimator for : The function is strictly convex on and by Jensen’s inequality, lemma A.7, we have , i.e. . The inequality is strict, since is not constant with probability one (by proposition A.3 it is a multiple of a chi-squared random variable). Thus, , the estimator is biased for and the Cramer–Rao bound, which only concerns unbiased estimators, cannot be compared to directly.
∎
- 1.
-
2.
If is known, the model consists of one parameter and by theorem 11.3 the Fisher information is . The lower right corner of is the diagonal element , which by definition is the variance of with fixed to the true value, i.e. for known shape. The diagonal element of the inverse answers a different question, asking about the estimation of when is unknown. Part (d) shows that this value is larger.
- 3.
-
4.
We have , and and thus the determinant of is . The lower right element of the inverse is given by
For a sample of size , the information matrix is , the inverse is and by theorem 14.5 we find the bound for an unbiased estimator for as claimed. The bound for known shape is and the ratio of the two is
where is as in (14.2). For we have and thus and the ratio is approximately : not knowing the shape increases the bound for the rate by a factor of approximately , as in section 14.5.
∎
-
1.
Since the two samples are independent, the log-likelihood is the sum of the two sample log-likelihoods:
with partial derivatives and . The second partial derivatives are , and , since each first partial derivative involves only one of the two parameters. None of these second derivatives are random, and thus proposition 14.3 gives , without any expectations.
-
2.
The difference is with and the inverse information matrix is . By theorem 14.5 any unbiased estimator for has variance at least
The estimator is unbiased for by lemma 2.2 and, since the two samples are independent, the variance of this estimator is . This is the bound. Thus, is an efficient estimator for the difference of the means.
-
3.
If we include as an additional parameter, the first derivative is
The mixed derivatives with respect to and and with respect to and are and , both with expectation zero, and the second derivative with respect to has expectation , as shown in example 14.6. Thus, the information matrix is
still diagonal, and thus is orthogonal to both means. The upper left block of the inverse is the inverse information matrix from part (b), and the Cramer–Rao bound for is still , and this bound is achieved by , whose definition does not depend on . For point estimation of the difference, not knowing the variance does not change anything. However, once we want to determine the accuracy of an estimate, the unknown variance starts to matter, since the bound depends on and thus needs to be estimated. This is the reason why we go from the normal distribution to the distribution in lectures 25 and 29.
∎
C.11 Consistency of the MLE
-
1.
For (R1) we have two different rates and thus two different distributions: for example, the means are different. For (R2) the support is for all . For (R3) we have and the derivatives w.r.t. are , and , all of which are continuous at . The derivatives of are of the form (polynomial in and ) times and thus are integrable over uniformly for in any compact subinterval of . We can therefore interchange the order of differentiation and integration, and is continuous and lies strictly between zero and infinity, so that (R3) holds. Finally, for (R4), the third derivative does not depend on and for with we have and thus . The constant bounds the third derivative and has finite expectation.
-
2.
The likelihood equation has only one solution , since with probability one. Since , the log-likelihood function is strictly concave and thus the root is the global maximum of the likelihood function, i.e. the MLE as in example 5.2. Since all conditions of theorem 16.3 are satisfied, we find that is consistent for .
- 3.
∎
-
1.
Since for , we have
and taking expectations with gives
-
2.
Using the rule for logarithms, we find
For to hold, we need to have equality in both applications of the inequality, i.e. we need and . Both of these conditions state . Thus we have for all , as stated in theorem 16.2.
∎
-
1.
Since all observations are at most , we have and thus is the same as . This is the case if and only if all observations are less than . Using independence we find
since . For , the probability of this event is zero. Thus and the MLE is consistent.
-
2.
Condition (R2) is violated, since the support depends on and (R3) is also violated: for fixed , the function has a jump from to at and thus is not differentiable there. The uniqueness of the MLE fails, since the likelihood equation has no roots: on the region where is positive, the derivative never equals zero. Since the proof of theorem 16.3 requires a differentiable likelihood and a root of the derivative to apply Rolle’s theorem, neither of these conditions are satisfied. Nevertheless, the statement is true, as can be seen by the direct computation in the first part of the question, or using the mean squared error from exercise 2.2.
∎
C.12 Asymptotic Normality of the MLE
- 1.
- 2.
-
3.
Substituting by in the asymptotic variance, we find the standard error and thus the approximate interval . Both of these quantities are undefined if all observations are equal, i.e. at the boundary of the domain discussed in section 17.5.
∎
-
1.
Using the result from the theorem, we find . Since the Poisson distribution has mean and variance equal to , this is the central limit theorem for .
-
2.
For we have and using the delta method we find
In exercise 5.7 we found and expanding we find that the bias is approximately , which is of order . Since it is multiplied by , the bias goes to zero and thus is of a smaller order than the spread and does not contribute in the limit: the estimator is biased for finite , but is asymptotically normal with mean .
∎
-
1.
For and , the event is the event , i.e. that all observations are less than . Thus we have
where we used the standard limit . The right-hand side is the probability that an exponential random variable with mean is greater than . Thus, converges in distribution to the corresponding exponential distribution.
-
2.
Since , the second factor converges in distribution to the result from the first part, and the first factor is a constant which converges to zero. Thus, by Slutsky’s theorem (theorem A.16), the product converges in distribution to . This is equivalent to convergence in probability to zero. Thus we cannot have a limit where , since the error of the MLE is of order and not of order . This is a manifestation of the super-efficiency from lecture 13, seen through the lens of the limiting distribution: the variance of the bias-corrected maximum is of order , smaller than allowed by the Cramer–Rao bound, and thus neither the bound nor theorem 17.1 apply, since the model violates (R2).
∎
C.13 The EM Algorithm
-
1.
Using the formula for the sum of a geometric series, we find
Since the event is contained in the event , we can use the definition of conditional probability to get
This shows that, given that the first trials were failures, the number of additional trials required until the first success is still geometrically distributed. The information about the number of failures so far does not change the distribution of the number of additional failures required. This is the discrete analogue of the statement for the exponential distribution in exercise 5.6. The shift by one is due to the fact that the support of the geometric distribution starts at instead of at .
-
2.
Given , the conditional probability weights of are
for . Substituting with we get
where the first sum is the total probability and the second sum is the mean of the geometric distribution. Using we get the claim. Alternatively we can get the same result from the first part of the solution: given , the value is geometrically distributed with parameter and thus has mean .
∎
-
1.
The log-likelihood function for the given data is
Taking derivatives we get the likelihood equation
This equation cannot be solved analytically for general densities.
-
2.
The density of the complete data is and . Thus we have
Using Bayes’ rule we find the conditional probability of given under as . Substituting the indicators using these responsibilities we get
where the constant term combines the terms and , which do not depend on . Setting the derivative equal to zero we find and the second derivative is negative for every , so is concave on and the stationary point is the global maximum.
-
3.
Using the notation for the denominator, we have
Summing over and comparing to the first part of the solution we find . A value is a fixed point of the iteration if and only if the left-hand side is zero at and since this is the case if and only if . Thus the fixed points of EM are the roots of the likelihood equation as given in proposition 19.3.
∎
At , the mixture density of observation is . Using the standard normal density values , , , and , we find the five densities to be , , , and . The corresponding logarithms are
and thus . At the densities are , , , and . The corresponding logarithms are
and thus . The log-likelihood has increased by as required by theorem 19.2. The largest contribution, , comes from observation : under this observation is three standard deviations away from the nearer mean and has tiny density, whereas under it is about one standard deviation away from the mean . The observations and contribute and , respectively, for the same reason. The observations and , which were always within one standard deviation of a mean, contribute little. ∎
-
1.
Since we know , the complete-data log-likelihood is the same as in example 19.4 and the terms and are now constants. The E step determines the responsibilities . Note that depends on only via the terms . The partial derivatives are as shown in the example, and setting them equal to zero gives the two mean updates of (19.2). The update for the weight is no longer needed.
-
2.
Since , we have and thus we can take derivatives of to get
and similarly we find . Setting both of these equal to zero gives and the corresponding equation for , stating that is mapped to itself by the updates of the first part. Thus, the likelihood equations and the fixed-point equations of the EM algorithm coincide as predicted by proposition 19.3, and we also find the identity of this proposition directly: the derivative of equals the derivative of at .
-
3.
As , the responsibilities converge to for all observations since . Thus, the update for is now independent of the initial value, and this is the MLE for the mean of a single normal sample in example 5.4, reached in one step. The update for is now the undefined ratio : a component with no weight has no observations to use to estimate its mean, and the mean of this component is not identifiable as discussed in section 19.5.
∎
C.14 Hypothesis Testing
-
1.
By equation (20.2) the power at is , where . Since is strictly increasing with , we have if and only if , i.e. if and only if
Since we have , and thus the right-hand side is positive; squaring both sides gives the stated condition.
-
2.
With and the bound is
and thus observations are needed.
-
3.
The bound is proportional to . Halving the difference multiplies the required sample size by four, to in the numerical example (the bound becomes ), and doubling has the same effect. Detecting a difference half as large costs four times as many observations.
∎
-
1.
Under the total is binomially distributed with parameters and . Thus, the test size is
-
2.
We can compute the binomial coefficients , and to find the sizes for , for and for . The smallest value with size less than or equal to is with size . Since can only take the values , the size can only take the eleven values computed in the first part of the question, and is not one of these values (the size jumps from to between and ).
-
3.
The power at is . This gives the values for , for and for . A coin which shows heads four times in five experiments is detected in fewer than of the experiments; for ten experiments the test can only detect large deviations from fairness with reasonable probability, and a failure to reject the hypothesis is not very meaningful.
-
4.
The -value is the probability of the event that the total under is larger than or equal to the observed value, i.e. the probability of . Since this probability is less than or equal to , we can reject at a significance level of . This agrees with the critical region from the second part of the question.
∎
-
1.
The maximum of the distribution is at most , if and only if all observations are at most . Since the observations are independent, for we have
This is also the integral of the density from example 8.10.
-
2.
For the size is , as found in the first part. Setting this equal to gives .
-
3.
For the power function is
This function increases from for to as . For , and we have , and for the power is .
-
4.
Under all observations are in the interval and thus we have : The second test has size and is a level test for every . The power function of this test is for . This function equals for in the numerical example. The difference between the two power functions is , and this is positive for every alternative, i.e. the first test is more powerful everywhere on , but the first test has type I error probability which the second test avoids. In the context of section 20.3, the first test is to be preferred: we have agreed on a level , and a test which does not use the level would be a waste of power. The power gain is at most , and decreases quickly as increases, whereas for the second test a rejection is certain if the observation is above , which is impossible under , and so we can reasonably assume that this is worth the small loss of power. The example shows that the choice of level is a matter for the statistician, and is not a mathematical necessity.
∎
-
1.
The observed value of the standardised statistic is and the one-sided -value is . Since we can reject at level , but since we cannot reject at level .
-
2.
The two-sided -value is . Again we reject at level and we cannot reject at level .
-
3.
The mean is a fixed, unknown number and not a random variable, so we cannot determine a probability for the statement . The -value is a probability about the data, assuming that . A correct interpretation is: if the true mean were , then a sample mean of at least would occur with probability , and thus the data are significant at level but not at level .
-
4.
Under the statistic is standard normally distributed. Since is continuous and strictly increasing, for we have
and thus the -value is uniformly distributed on under . In particular we have , which is the statement that the test has size .
∎
C.15 The Neyman–Pearson Lemma
-
1.
The joint probability weights are with . Thus the likelihood ratio is
Since , this function is strictly increasing as a function of and the inequality is equivalent to with . By theorem 22.2 the most powerful test at this level is the one which rejects when . The size of this test, , can be found using the Poisson distribution with mean , which contains the parameters and only. For a given size the threshold is the same for all . By proposition 22.4 the test is uniformly most powerful against .
-
2.
Under , the sum is Poisson-distributed with mean and thus we have , whereas . The smallest acceptable threshold for the test is : we can reject and the size of this test is .
-
3.
Under , the sum is Poisson-distributed with mean and thus the power of the test is .
-
4.
The size of the test is . This value can only take the values listed for and the jumps between these values are from to , i.e. to . The value is not one of the possible values. The test from part (b) has size and the lemma shows that this is the most powerful test at level . The test is also a test of level , but the lemma does not show that this is the most powerful test among all tests with level . As the remark in section 22.2 points out, a randomised test could have size exactly and higher power.
∎
-
1.
The joint probability weights are with . Under , these weights equal for all . Thus we have
as claimed. For , we have and thus is strictly increasing in . The inequality is equivalent to for some . By theorem 22.2, the most powerful test is to reject when , where is determined by the distribution of under , i.e. the binomial distribution with parameters and . Since this distribution does not depend on , the same region of is optimal for all alternatives and by proposition 22.4 the test from exercise 20.2 is uniformly most powerful against .
-
2.
Since and , the largest critical region which does not exceed size is . This region has size . The power of this test at is . For , the corresponding region from exercise 20.2 has smaller size but the power at is . By doubling the sample size, we can use a test with size closer to the nominal value and nearly triple the power at this alternative.
-
3.
For , we have and thus is strictly decreasing in . The inequality is now equivalent to and the most powerful test will be a test which rejects for small values of . For , with the same size , the test will reject for , by the symmetry of the binomial distribution with parameter . Thus, the two one-sided problems will result in tests which reject on the opposite tails of the distribution of and there can be no single critical region which satisfies both tests.
∎
-
1.
For we have
and since there are three cases. If , both densities are positive and
If , we have , and thus . If , both densities vanish and equality holds. Consequently
-
2.
Since , the size of is . The region contains the set of part (a), because , and it is contained in the union of the two sets of part (a), which is the whole sample space. Thus satisfies the condition of theorem 22.2 with and by the first statement of the theorem it is most powerful at level . The power of this test is
The threshold does not depend on and the argument from above applies to all with the corresponding and thus, by proposition 22.4, the test is uniformly most powerful against .
-
3.
The two parts of are disjoint and thus the size of this region is . The power of this test is
This is the same power as for . Thus both regions are most powerful at level and, by the second statement of theorem 22.2, they can only differ on the boundary set. And indeed they do: the two regions differ on the sets and , both of which are contained in where . The argument in the uniqueness comment in section 22.2 used the fact that the likelihood ratio has a continuous distribution, but in this case this is not true: on the set , where an event with probability occurs under , the ratio is constant and thus the boundary set has positive probability, despite the model having densities.
∎
Since is sufficient, we can use the factorisation theorem 8.2 to find functions and such that
for all and all parameter values . Now let us choose an with . Then neither of the two factors in the expression for can be zero and we get
Since we can cancel the factor on both sides, the right-hand side only depends on via . This completes the proof. ∎
C.16 Likelihood-Ratio Tests
Here we solve the exercises at the end of lecture 23.
-
1.
By example 5.4 the log-likelihood
is maximised over at and . Substituting these values, the sum of squares becomes and the second term becomes , and thus the maximised log-likelihood is
- 2.
-
3.
By (23.1) and the two previous parts, the terms cancel and we get
Since , we have , which gives the stated formula. The function is strictly increasing, and thus the event is the event for a suitable : the generalised likelihood-ratio test rejects for large values of .
-
4.
The full parameter has coordinates, and fixes one of them, , leaving free, so that . The normal model satisfies the regularity conditions, and thus theorem 23.3 gives under . The approximate test rejects when , i.e. when . For this is , or , whereas the exact test rejects when . The approximate test rejects more often than the exact one, and its true size is instead of the nominal : the chi-squared approximation is optimistic for small , because the distribution has heavier tails than the normal limit which the approximation assumes. For large the two tests merge, since by and the distribution tends to the standard normal one; for the two thresholds are and . Whenever the exact distribution is available, as here, it is the one to use, and lecture 29 does so.
∎
-
1.
Taking logarithms of the probability weights gives
where the first term does not depend on . To maximise subject to we consider the Lagrangian . Setting its partial derivative with respect to to zero gives , i.e. for every , and summing over and using the constraint we find , so that and . This stationary point is the maximum, because is a concave function of and the constraint set is convex. If some is zero, the corresponding lies on the boundary of the parameter space; the formula still gives the maximum, since any probability assigned to an unobserved outcome only lowers the likelihood.
-
2.
Since is simple, the restricted maximiser is , and by (23.1) we get
where the combinatorial constant cancels. A term with contributes nothing to either sum, which is the convention .
-
3.
The parameter space is described by the free coordinates , with determined by the constraint, and thus has dimension . The null hypothesis fixes all of them, so that . The counts are the sums of i.i.d. indicator vectors, one per repetition, whose distribution is a full-rank exponential family with support not depending on , so that the regularity conditions hold at every interior point . Theorem 23.3 therefore gives under .
-
4.
Under the expected counts are for every face. The terms are
with sum , and thus . Since , we do not reject at level ; the approximate -value is . The observed counts are entirely consistent with a fair die.
-
5.
Write , so that and . Using we get
since and . Summing over , the linear terms add up to , and doubling the result gives
up to terms of third order in the relative deviations . For the die the deviations from are , and Pearson’s statistic is , close to the value of ; the two statistics agree to the extent that the relative deviations are small, which for these data they are.
∎
-
1.
By example 4.2 the likelihood is , so that , and by example 5.3 it is maximised over at when . Since is simple, the restricted maximiser is , and (23.1) gives
For or the supremum of over is approached at the boundary, where one of the two terms is absent, and the convention gives the correct value.
- 2.
-
3.
We have and , and thus
Since , we reject at level . The approximate -value is , just below , so that the evidence against a fair coin is real but not strong.
-
4.
The map is a strictly increasing bijection from onto with inverse , and the likelihood in the new parametrisation is . The null hypothesis is the same set of distributions as with , and thus the numerator of the ratio is in both cases. Since the supremum is attained at the MLE by the first part, and is a bijection from onto , the two suprema and run over the same set of values, and thus the denominator is . Both ratios are therefore equal, and so are the two statistics . This is the invariance noted in section 23.1; by contrast the standard errors of exercise 17.1 do depend on the parametrisation.
∎
C.17 Confidence Intervals
-
1.
The function is a function of the data and the parameter, and from exercise 13.1 we know that under the distribution of is for all , independent of . Thus, is a pivot. Using and we find and since the function is increasing in , the two inequalities can be solved for directly:
Thus, the interval
is a confidence interval for with level and the probability of this interval covering the correct value is exactly .
-
2.
We have and and thus the interval is
-
3.
The MLE for the rate is and the standard error of the MLE is . Thus the Wald interval is , which is the interval . The exact interval is asymmetric about the estimate, extending further to the right than to the left. This is caused by the right skew of the distribution of for small exponential samples; the Wald interval, being symmetric, extends to lower values than the exact interval does. At the normal approximation used in the Wald interval is rough, as noted in section 17.5, and the exact interval should be used instead.
∎
-
1.
From example 8.10 we know that the maximum has density for . Thus, for we have
The distribution function of does not depend on and thus is a pivot.
-
2.
Since always holds, we have
where we used the first part of the solution. Thus, is a confidence interval with level . For and we have and thus the interval is .
-
3.
The condition for coverage is . This condition can be used to find as a function of and thus, by choosing the interval , we can find the length of the interval. Taking derivatives we find and thus
since . The length of the interval decreases as increases and thus is smallest for the largest possible value , where . This is the interval from the previous part of the solution. The reason for this result is that the density of the pivot is increasing on and thus the probability is concentrated on the shortest possible interval which is placed at the upper end of the density.
∎
-
1.
The standard error is , and the Wald interval is . For the data we have and , and the interval is , which is .
-
2.
In exercise 17.2 we found , and replacing by in the variance gives the standard error , in agreement with the delta-method formula for . The approximate interval is . For the data we have and , and the interval is , which is .
-
3.
Write for the interval of the first part. Since is decreasing, the event is the same as the event , and thus the interval covers exactly when covers . Its coverage probability is that of the Wald interval, which converges to by proposition 25.7. For the data we get . The two intervals differ because is not linear: the delta-method interval is symmetric about , while the transformed interval is shifted towards larger values and is asymmetric about . Both have asymptotic coverage , and for large the two nearly coincide, since is close to linear over a short interval. A practical advantage of the transformed interval is that it always lies inside , whereas the delta-method interval can extend below zero when is small.
∎
-
1.
Under , the test statistic is the pivot from example 25.4 and has distribution, independent of the value of . Since the distribution is symmetric, we have
and the test has size .
- 2.
-
3.
The -interval for the given data is . By duality, is rejected at level if and only if is outside this interval. Thus, is rejected since , and is not rejected since is inside the interval.
-
4.
The one-sided test does not reject if and only if , i.e. if . Thus, the set of values not rejected is the one-sided interval , whose left-hand boundary is a lower confidence bound for with level . Under , the test statistic is distributed, and thus the test has size and the bound has coverage probability exactly , by theorem 25.6.
∎
C.18 Bayesian Inference: Priors and Posteriors
-
1.
The prior mean is and since we have and . From table A.2 we find that the variance of the distribution is
and thus the prior standard deviation is approximately .
-
2.
Let denote the prior mean and the equivalent sample size. Then we have and . Using the formula for the variance from the table, we get
Given and a variance of we find that we need , i.e. and thus we have and . The equivalent sample size is : the smaller spread corresponds to a prior which gives twice the number of observations.
- 3.
∎
-
1.
On the interval we have and thus there. Since , the integral of the prior over is infinite and the prior is improper.
-
2.
Consider the kernel for real and . Near the factor is bounded between two positive constants and thus the kernel is integrable near if and only if is, which is equivalent to . By symmetry the kernel is integrable near if and only if and thus, since it is bounded away from the endpoints, the integral over is finite if and only if and , in which case it equals . For the formal posterior we have and and thus the posterior is proper if and only if and , i.e. if and only if . In this case it is the distribution.
- 3.
-
4.
For the likelihood has its maximum at and does not vanish there, while the prior has a non-integrable singularity at . Since the data does not contain any successes, the infinite mass of the prior near zero cannot be compensated and thus the integral of is infinite: there is no posterior distribution in this case. In informal language, the improper prior suggests that is very close to either or and a data set with no successes cannot exclude the first of these two possibilities.
∎
- 1.
-
2.
Since the prior density, as a function of , is proportional to and since we can group the terms in the exponent of the posterior density as powers of , we get
where we have omitted the terms which do not contain . Using the abbreviations and we can write the exponent as and completing the square we find
where we have omitted the last term which does not contain . Thus we have and this is the kernel of the density and, since the posterior is a density, this is the corresponding normal distribution. If we multiply out we get the weighted form of the posterior given in proposition 26.6.
-
3.
As with fixed, the prior precision goes to zero and thus and : the posterior converges to , i.e. to the posterior under the flat prior discussed in section 26.4. As with fixed, the weight of the prior mean goes to zero and thus and goes to zero: the posterior is now concentrated around and the prior is washed out.
∎
-
1.
Under , the total number of successes is binomially distributed with and . Since is an affine function of , we have
which equals zero if and only if , the prior mean, and
The bias is of order and the variance is of order , both tending to zero as . Thus, by corollary 2.11, the estimator is consistent, although for fixed it is biased for all .
-
2.
For and we have and thus, by theorem 2.7,
At the bias is zero and we have compared to . Thus, the Bayes estimator is better. At we have compared to and thus is better. Neither estimator has smaller mean squared error for all : the Bayes estimator is better in the middle of the parameter range, where the prior mean is close to the truth, and the reduced variance does not compensate for the bias introduced by the shrinkage towards at the ends of the range.
∎
For the first part of the question we can take derivatives. The log-likelihood function for a single observation is and thus the score function is
Since , with density , for all values of , the Fisher information is
This does not depend on .
For the second part of the question we have and thus the score function is
where . Since has density for all , we find
does not depend on .
Finally, the model is the scale model with base density the standard normal density and with and thus . In this case we have and using the substitutions and for the standard normal density we find
Thus , as claimed. ∎
C.19 Credible Intervals and Bayesian Testing
-
1.
Since the posterior density is monotone on each half-line and integrates to one, it converges to zero as . The condition excludes , where contains at most the point and has probability zero, and thus . On the density is continuous and strictly increasing from to and thus there is exactly one with and for if and only if . Similarly, there is exactly one with and the density is at least on exactly on . Thus we have with .
-
2.
Let and with . By symmetry we have and since the density is strictly decreasing on we find . The two tails and have equal posterior probability by symmetry and together they have probability , i.e. each tail has probability . Thus we have and , i.e. we have found the equal-tailed interval.
-
3.
The set is for the interval with , obtained by using the same argument as in the first part, but only for the decreasing branch of the density, and using the condition to get . The equal-tailed interval is then obtained from by removing the piece and adding the piece , both of which have posterior probability . The density is at least on the first piece and at most on the second piece, where since the density is strictly decreasing. Writing and for the lengths of the two pieces we find and and thus : the added piece is longer than the removed piece and the equal-tailed interval is longer than the HPD interval.
∎
-
1.
The equal-tailed credible interval is , with posterior probability cut off in each tail. The midpoint of this interval is , which is slightly below the posterior mean .
-
2.
The posterior density is proportional to on . Setting the derivative of the logarithm of the posterior density, i.e. , equal to zero, gives the maximum of the density at and thus the mode is at . Since the mode is above the mean , the posterior distribution is skewed.
-
3.
By definition 28.2 of the HPD interval, the HPD interval contains the values with the highest posterior density and thus is centred around the mode, so that it lies to the right of the equal-tailed interval. The HPD interval is the shortest interval which has posterior probability , in this case of length against . As increases, the posterior distribution approaches a symmetric, normal distribution and the mode and mean move together. In the limit, for a symmetric, unimodal posterior, both intervals coincide; this is a result from exercise 28.1.
∎
-
1.
Since the prior is flat, we have and using and from example 26.5 we find the posterior . Since , the density of this posterior is for .
-
2.
We can find the posterior probability of by integrating the density:
Thus we have and the posterior odds for against are .
-
3.
The data have multiplied the odds by . Using Bayes’ theorem for the events and , we find that the posterior probability of is times the average of the likelihoods over , with respect to the prior conditioned on , divided by the marginal density of the data. The factor by which the odds change is the ratio of these two averaged likelihoods. In this case, the likelihood is and the conditional priors are uniform on and on , and thus the factor is
This agrees with the result from the previous part. Note that this is not the ratio of likelihoods at two values, since each hypothesis includes a range of parameter values, and these need to be weighted by the prior before the two hypotheses can be compared. Only for two simple hypotheses do the averages coincide with single likelihood values, as in case (28.3).
-
4.
The bounds and satisfy and , i.e. we have
These are polynomial equations of degree four and while they can be solved explicitly, the solutions are cumbersome and not very instructive; for a general Beta posterior, the bounds will not be algebraic. The quantiles are best found numerically, for example using the R function qbeta to get and and thus the equal-tailed credible interval is .
∎
-
1.
Using the formula from example 28.7, with , , , and , we find that the Bayes factor in favour of is
The data give evidence in favour of over by a factor of .
-
2.
The prior odds are and thus the posterior odds are and . The data have changed the prior opinion, which was three to one in favour of , to a posterior opinion of approximately four to one against it.
-
3.
The -value is the probability, under , of obtaining a sample mean of at least . This value refers to the data, includes outcomes more extreme than the one observed and makes no reference to any alternative. In contrast, the posterior probability is a probability for the hypothesis; it uses only the observed data, compares the likelihoods at exactly and and takes into account the prior odds, which were chosen in favour of . Even if the prior odds had been , the posterior probability would have been , or six times the -value. Both numbers show that the data are evidence against , but in different ways. Neither number is a rewording of the other.
-
4.
The posterior odds are the prior odds times , and the posterior odds are greater than if and only if the prior odds are greater than . Only a prior which assumed to be more likely than , by a factor of more than twelve, would result in being the more likely hypothesis after these data are observed.
∎
- 1.
-
2.
Write and for the standardised distance between the prior mean and the true value. The interval covers if and only if . Subtracting the mean and dividing by the standard deviation turns into a standard normal variable, and the bounds into . Thus the coverage probability is
It depends on only through , and it is symmetric in .
-
3.
Here and , and thus , , and . At we have and the coverage probability is , above the nominal . At we have and , and the coverage probability is , below the nominal value. As we have , both arguments of tend to the same infinity and the coverage probability tends to zero: the interval is pulled towards the prior mean by a fixed fraction of the distance, and once that pull exceeds the half-width the interval almost never contains the truth.
-
4.
It is not a confidence interval with level , because definition 25.1 requires coverage probability at least for every , and the previous part exhibits values of where it is smaller. The credible interval trades the uniform guarantee for better behaviour near the prior mean, which is where a well-chosen prior expects the truth to be. As with fixed we have , and thus and , so that the coverage probability tends to for every . This is the large-sample agreement of section 28.3 seen from the frequentist side.
∎
C.20 Standard Tests Revisited
-
1.
We have the sum of observations and thus . The sum of squared deviations from equals and thus we have and . The observed value of the test statistic is
-
2.
From proposition 29.1 we know that the test against rejects if . For level the critical value is and for level the critical value is ; since lies below both of these critical values, the hypothesis is rejected at both levels. The -value is given by for , by symmetry, and since is between and , the value is between and .
-
3.
The two-sided test rejects if . For level we have and thus is rejected. For level we have and thus is not rejected. The two-sided -value is , which is twice the value of the one-sided test, so it is between and .
-
4.
The interval is . This interval is . The interval contains the value and by theorem 25.6 this is equivalent to the non-rejection of by the two-sided test with level in the previous part. The interval does not contain and thus the rejection at level is consistent with the results of the other parts.
∎
-
1.
The value of the test statistic (29.2) is . We can compare this value to quantiles of the distribution: at level the critical value is and since we can reject , whereas at level the critical value is and since we cannot reject . The -value, for , is between and , since is between and .
-
2.
The two-sided test with level does not reject, if
Since , we reject .
-
3.
Let . Assuming the true variance is , the value is distributed by proposition A.3 and thus . We have
As increases, the threshold decreases and the probability of exceeding a smaller number increases, i.e. increases as a function of . By definition 20.5, the size of the test for the composite null hypothesis is , where the bound is achieved at the boundary.
-
4.
The two-sided test with level rejects large values of only if . Since this is a larger threshold than the of the one-sided test, the power of the two-sided test against the alternative is smaller than the power of the one-sided test. An observed value between the two thresholds is rejected by the one-sided test but not by the two-sided test. The two-sided test has power above against the alternative , where the one-sided test has probability of rejection below . For the given data, both tests reject the hypothesis at level , but at level the one-sided test, with threshold , still rejects, whereas the two-sided test compares to and does not reject.
∎
-
1.
We can describe the differences using i.i.d. , where both parameters are unknown. We test against , since we expect the reaction time to be smaller after training. The differences are and the sum of these differences is , i.e. we have . The sum of squared deviations from is , and thus we have and . The test statistic from proposition 29.1, for , is
where is the number of degrees of freedom. Since we can reject the hypothesis at level and since we can also reject the hypothesis at level : the data show a reduction in reaction time.
-
2.
The -interval from example 25.4 for is , which is .
-
3.
The row means are and , the pooled variance is and thus we have . The test statistic (29.3) is
where is the number of degrees of freedom. Since the colleague does not reject the hypothesis at level . The two-sample test is wrong, because the two measurements for one subject are not independent (a slow person is likely to be slow after training, too), whereas the two-sample test assumes that the twelve observations are independent. The difference in opinion is caused by the variability of the data. Between subjects the standard deviation of the reaction times is about . The two-sample test compares the mean difference to this spread. Inside subjects the differences have standard deviation only . The paired test compares to this smaller spread. By taking differences we remove the effect of the subjects from the comparison, and this is the idea of the paired test.
∎
- 1.
-
2.
From theorem 25.6 we know that the two-sided test from proposition 29.3 rejects at level if and only if is outside the confidence interval from example 25.5. This is the inverse of the test, so is rejected and is not. Looking at the values directly, for the test statistic corresponding to is , which is greater than , so is rejected. For the test statistic falls into the interval and thus is not rejected.
-
3.
For the test statistic is and, since , the one-sided test does not reject at level . The two-sided interval is the inverse of the two-sided test with level , i.e. the upper tail of the two-sided test has probability . If we decided a one-sided test by checking whether is less than the lower bound of the interval, we would obtain the one-sided test at level , not at level . Thus, the one-sided test with is not the same as the two-sided test. The confidence set corresponding to the one-sided test with is, as in exercise 25.4, the one-sided interval , and is contained in this interval, as expected.
∎
C.21 The Exponential Model: From Start to Finish
Here we solve the exercises at the end of lecture 31.
-
1.
The log-likelihood function is , with , and the MLE is . Using and we find
Using we have and , and thus .
-
2.
Let for . Then and . Since this value is negative for and positive for , must be the unique minimum of and for . Since , we have , with equality if and only if , i.e. if and only if .
-
3.
For we have and . This value is less than , so is not rejected at level . For we have and . This value is greater than , so is rejected: the data are not compatible with a mean lifetime of one year.
This completes the solution. ∎
-
1.
From example 11.8 we know that . Thus, the Jeffreys prior is . Since , the prior has infinite integral and thus is improper.
-
2.
The posterior density is proportional to . This is the Gamma density up to the constant . Since , this constant is finite, so the posterior is proper and corresponds to the Gamma distribution.
- 3.
-
4.
For we have and thus the equal-tailed credible interval is . These are the same numbers as for the exact confidence interval. The confidence interval covers the fixed value in of repeated samples. The credible interval, on the other hand, covers the random variable with posterior probability given the observed data.
This completes the solution. ∎
- (a)
-
(b)
By part (a) and the linearity of expectation, for all , and thus is unbiased. The same substitution gives
and thus
and . Since is unbiased, as , and thus the mean squared error does not tend to zero. This alone does not disprove consistency, since theorem 2.10 is only a sufficient condition; but directly, for we have
and thus does not converge to in probability: is not consistent in the sense of definition 2.8. The reason is that uses only the smallest observation, which carries little information about the upper endpoint .
- (c)
-
(d)
Conditionally on , the other observations are i.i.d. uniform on , and is their minimum. By part (a) with in place of and in place of , the minimum of i.i.d. uniform observations on has mean , and thus
The Rao–Blackwell theorem, theorem 10.1, applies because is sufficient by part (c) and is unbiased with finite variance by part (b). It guarantees that does not depend on , which the formula confirms, that is unbiased for , and that .
-
(e)
By proposition 10.5 the maximum is complete, and it is sufficient by part (c). Since is a function of and unbiased for , the Lehmann–Scheffe theorem, theorem 10.4, shows that is the UMVUE for . Its variance is smaller than by a factor , and it lies below the Cramer–Rao value . There is no contradiction, because theorem 13.2 does not apply to this model: its conditions, section 11.1, require the support of the density not to depend on , and here the support is . The identity on which the proof rests fails, because differentiating with respect to picks up a term from the moving boundary of the integral. The value is therefore not a bound at all for this model, and estimators based on the maximum are super-efficient in the sense of example 13.6.
∎
-
(a)
The tray comes from supplier A with probability , and given the supplier the count is binomial, and thus by the law of total probability
a two-component mixture with known components. For independent trays the observed-data log-likelihood is
The likelihood equation is, after clearing denominators, a polynomial equation of degree in , and the sum inside the logarithm prevents the simplification which made the examples of lecture 5 solvable. The complete data are the pairs with and , and thus
which is linear in the indicators.
- (b)
-
(c)
Differentiating with the held fixed gives
which vanishes at , is positive to the left of this value and negative to the right. Thus is the maximiser; equivalently, the second derivative is negative for every , so is concave on and is its global maximiser. The update is the complete-data MLE of example 5.3 with the unobserved indicators replaced by responsibilities.
-
(d)
With the weights cancel and . From the table we get, to four decimal places,
and thus and . The responsibilities are all close to or . Since the two component distributions barely overlap, each count identifies its supplier almost with certainty, and the trays with , and germinated seeds are assigned to supplier A, the trays with and to supplier B. The tray with seeds is the least certain, at . Further iterations move only slightly, to the fixed point ; the observed-data log-likelihood rises from at to at , as theorem 19.2 requires.
-
(e)
With unknown, can no longer be dropped from , which now contains the term
Setting the derivative of this term with respect to to zero gives
the responsibility-weighted proportion of germinated seeds among the trays attributed to supplier A, with the responsibilities now computed from . For the data of part (d) and the responsibilities computed there this gives . The ascent property, theorem 19.2, states that the observed-data log-likelihood never decreases along the EM sequence, for all . It guarantees that the sequence of log-likelihood values converges, and by proposition 19.3 a point at which the iteration comes to rest is a stationary point of ; it does not guarantee that this point is the global maximum, since a poor starting value can lead to a local maximum, and it says nothing about the speed of convergence.
∎