Lecture 7 Exponential Families

In lectures 4 and 5 we repeatedly collapsed products of nn densities to expressions which summarised the data using only one or two sums. This lecture explains the reason for this coincidence: almost all families of models considered here have the same algebraic shape, and models of this shape are called exponential families. This shape allows us to define the data summary used in the likelihood, the moments of this summary, and the smoothness required for the later theory, and many of the results in the following weeks are almost automatic for this class.

7.1 Recognising the Shape

We start by considering a familiar distribution and then rewrite its probability weights so that the data and the parameter are as far apart as possible. For the Bernoulli distribution with x∈{0,1}x\in\{0,1\} we can write

f⁢(x;θ)\displaystyle f(x;\theta)
=θx⁢(1−θ)1−x=exp⁡(x⁢log⁡θ+(1−x)⁢log⁡(1−θ))\displaystyle=\theta^{x}(1-\theta)^{1-x}=\exp\bigl{(}x\log\theta+(1-x)\log(1-% \theta)\bigr{)}
=exp⁡(x⁢log⁡θ1−θ+log⁡(1−θ)).\displaystyle=\exp\Bigl{(}x\log\frac{\theta}{1-\theta}+\log(1-\theta)\Bigr{)}.

Here we have written the weights as the exponential of the product of one function of θ\theta and one function of xx, minus a function of only θ\theta. The Poisson weights and the exponential density can be written in a similar way: we have θx=ex⁢log⁡θ\theta^{x}=e^{x\log\theta} and θ⁢e−θ⁢x=e−θ⁢x+log⁡θ\theta e^{-\theta x}=e^{-\theta x+\log\theta}, up to a factor depending on xx. We will give a name to this structure.

Definition 7.1.

A parametric model (f⁢(x;θ)|θ∈Θ)\bigl{(}f(x;\theta)\!\mathrel{\big{|}}\!\theta\in\Theta\bigr{)} with Θ⊆ℝ\Theta\subseteq\mathbb{R} is a one-parameter exponential family, if the density or the probability weights can be written as

f⁢(x;θ)=h⁢(x)⁢exp⁡(η⁢(θ)⁢T⁢(x)−B⁢(θ))f(x;\theta)=h(x)\exp\bigl{(}\eta(\theta)\,T(x)-B(\theta)\bigr{)}

for all xx and all θ∈Θ\theta\in\Theta, where h≥0h\geq 0 and TT are functions of xx alone and η\eta and BB are functions of θ\theta alone. In particular, the support {x∣f⁢(x;θ)>0}\{x\mid f(x;\theta)>0\} is {x∣h⁢(x)>0}\{x\mid h(x)>0\} and does not depend on θ\theta. The statistic T⁢(X)T(X) is called the natural statistic and η⁢(θ)\eta(\theta) is called the natural parameter of the family.

Since the exponential is always positive, f⁢(x;θ)f(x;\theta) is positive if and only if h⁢(x)h(x) is. This proves the statement about the support. The representation is not unique, since we can replace η\eta by c⁢ηc\eta and TT by T/cT/c for any constant c≠0c\neq 0. Finally, since the density must integrate (or sum) to one, the function BB is determined by the other three functions: we have

equation (7.1) (7.1)
B⁢(θ)=log⁢∫h⁢(x)⁢exp⁡(η⁢(θ)⁢T⁢(x))⁢dx.B(\theta)=\log\int h(x)\exp\bigl{(}\eta(\theta)\,T(x)\bigr{)}\,\mathrm{d}x.

These rewritings give the following examples.

Example 7.2.

The Bernoulli distribution with θ∈(0,1)\theta\in(0,1) is a one-parameter exponential family with h⁢(x)=1h(x)=1 for x∈{0,1}x\in\{0,1\}, natural statistic T⁢(x)=xT(x)=x, natural parameter η⁢(θ)=log⁡(θ/(1−θ))\eta(\theta)=\log\bigl{(}\theta/(1-\theta)\bigr{)}, the log-odds of success, and B⁢(θ)=−log⁡(1−θ)B(\theta)=-\log(1-\theta).

Example 7.3.

The Poisson distribution with θ>0\theta>0 is a one-parameter exponential family with h⁢(x)=1/x!h(x)=1/x! for x∈{0,1,2,…}x\in\{0,1,2,\dots\}, natural statistic T⁢(x)=xT(x)=x, natural parameter η⁢(θ)=log⁡θ\eta(\theta)=\log\theta and B⁢(θ)=θB(\theta)=\theta.

Example 7.4.

The exponential distribution with rate θ>0\theta>0 is a one-parameter exponential family with h⁢(x)=1h(x)=1 for x≥0x\geq 0 and h⁢(x)=0h(x)=0 otherwise, natural statistic T⁢(x)=xT(x)=x, natural parameter η⁢(θ)=−θ\eta(\theta)=-\theta and B⁢(θ)=−log⁡θB(\theta)=-\log\theta.

The normal distribution is also an exponential family, but since both parameters are unknown, we need to use the two-parameter version of the definition from section 7.3.

Remark.

The uniform distribution on [0,θ][0,\theta] is not an exponential family, since the support of the distribution changes with the parameter. In contrast, the support of an exponential family is always fixed. This is the same reason why the likelihood equation was useless in example 5.5. Similarly, the tt location model from section 5.3 is not an exponential family, despite the fact that its support is all of ℝ\mathbb{R}; we do not prove this.

The shape survives the passage from one observation to a sample.

Lemma 7.5.

Let X1,…,XnX_{1},\dots,X_{n} be i.i.d. from the one-parameter exponential family in definition 7.1. Then the joint density of the sample is

f⁢(x1,…,xn;θ)=∏i=1nh⁢(xi)⁢exp⁡(η⁢(θ)⁢∑i=1nT⁢(xi)−n⁢B⁢(θ)),f(x_{1},\dots,x_{n};\theta)=\prod_{i=1}^{n}h(x_{i})\exp\Bigl{(}\eta(\theta)% \sum_{i=1}^{n}T(x_{i})-nB(\theta)\Bigr{)},

which is of the form of definition 7.1 with the same natural parameter η⁢(θ)\eta(\theta), natural statistic Tn=∑i=1nT⁢(Xi)T_{n}=\sum_{i=1}^{n}T(X_{i}), the function ∏ih⁢(xi)\prod_{i}h(x_{i}) instead of hh and n⁢B⁢(θ)nB(\theta) instead of B⁢(θ)B(\theta).

Proof.

Since the observations are independent, the joint density is the product ∏i=1nh⁢(xi)⁢exp⁡(η⁢(θ)⁢T⁢(xi)−B⁢(θ))\prod_{i=1}^{n}h(x_{i})\exp\bigl{(}\eta(\theta)\,T(x_{i})-B(\theta)\bigr{)} of the marginal densities. Combining the nn exponents using the rule ea⁢eb=ea+be^{a}e^{b}=e^{a+b} gives the required form. This completes the proof. ∎

Read as a function of θ\theta, the joint density is the likelihood, and thus the log-likelihood of an exponential family sample is

equation (7.2) (7.2)
ℓ⁢(θ)=∑i=1nlog⁡h⁢(xi)+η⁢(θ)⁢Tn−n⁢B⁢(θ).\ell(\theta)=\sum_{i=1}^{n}\log h(x_{i})+\eta(\theta)\,T_{n}-nB(\theta).

The first term does not involve θ\theta and thus, except for the term involving nn, the likelihood depends on the data only via the single number Tn=∑iT⁢(xi)T_{n}=\sum_{i}T(x_{i}). This is an extension of example 4.3 to the whole class; in lecture 8 we will see in which sense the value TnT_{n} contains all the information about θ\theta.

7.2 The Natural Parametrisation and the Cumulant Function

Since the parameter enters the density only via η⁢(θ)\eta(\theta), it is often convenient to use η\eta as the parameter. In this case, the normalising function (7.1) could be given a name, since its derivatives are the moments of the natural statistic.

Definition 7.6.

Let hh and TT be as in definition 7.1. Then the natural parametrisation of the family is given by

f⁢(x;η)=h⁢(x)⁢exp⁡(η⁢T⁢(x)−K⁢(η)),η∈𝒩︀,f(x;\eta)=h(x)\exp\bigl{(}\eta\,T(x)-K(\eta)\bigr{)},\qquad\eta\in\mathcal{N},

where the cumulant function (also called the log-partition function) is given by

K⁢(η)=log⁢∫h⁢(x)⁢eη⁢T⁢(x)⁢dx,K(\eta)=\log\int h(x)\,e^{\eta T(x)}\,\mathrm{d}x,

where the sum replaces the integral for discrete models, and the natural parameter space is the set 𝒩︀={η∈ℝ∣K⁢(η)<∞}\mathcal{N}=\{\eta\in\mathbb{R}\mid K(\eta)<\infty\} of all η\eta such that the integral is finite.

Since KK makes f⁢(x;η)f(x;\eta) integrate to one, the formula defines a probability distribution for every η∈𝒩︀\eta\in\mathcal{N}, whether or not η\eta is of the form η⁢(θ)\eta(\theta); comparing with (7.1) shows that η⁢(θ)∈𝒩︀\eta(\theta)\in\mathcal{N} and B⁢(θ)=K⁢(η⁢(θ))B(\theta)=K\bigl{(}\eta(\theta)\bigr{)}, so that the original family sits inside this larger one. For the three examples of section 7.1 we find K⁢(η)=log⁡(1+eη)K(\eta)=\log(1+e^{\eta}) with 𝒩︀=ℝ\mathcal{N}=\mathbb{R} for the Bernoulli distribution, K⁢(η)=eηK(\eta)=e^{\eta} with 𝒩︀=ℝ\mathcal{N}=\mathbb{R} for the Poisson distribution and K⁢(η)=−log⁡(−η)K(\eta)=-\log(-\eta) with 𝒩︀=(−∞,0)\mathcal{N}=(-\infty,0) for the exponential distribution. In all three cases 𝒩︀\mathcal{N} is an open interval and θ↦η⁢(θ)\theta\mapsto\eta(\theta) is a bijection onto 𝒩︀\mathcal{N}, making the natural parametrisation a relabelling in the sense of theorem 5.6.

The first two derivatives of KK give the mean and the variance of T⁢(X)T(X). The name of KK is derived from the fact that s↦K⁢(η+s)−K⁢(η)s\mapsto K(\eta+s)-K(\eta) is the cumulant generating function of T⁢(X)T(X), as given in exercise 7.6.

Theorem 7.7.

Let η\eta be an interior point of the natural parameter space 𝒩︀\mathcal{N}. Then, in the natural parametrisation,

𝔼η⁢(T⁢(X))=K′⁢(η)andVarη(T⁢(X))=K′′⁢(η).\mathbb{E}_{\eta}\bigl{(}T(X)\bigr{)}=K^{\prime}(\eta)\qquad\text{and}\qquad% \mathop{\mathrm{Var}}\nolimits_{\eta}\bigl{(}T(X)\bigr{)}=K^{\prime\prime}(% \eta).
Proof.

Set Z⁢(η)=∫h⁢(x)⁢eη⁢T⁢(x)⁢dxZ(\eta)=\int h(x)e^{\eta T(x)}\,\mathrm{d}x, so that K=log⁡ZK=\log Z. Since h≥0h\geq 0 is not identically zero and eη⁢T⁢(x)>0e^{\eta T(x)}>0, we have Z⁢(η)>0Z(\eta)>0 and thus the division in the proof is justified. Since η\eta is in the interior of 𝒩︀\mathcal{N}, we can take derivatives inside the integral as many times as required; this is a result from analysis which we use without proof. Taking derivatives once, we get

Z′⁢(η)=∫T⁢(x)⁢h⁢(x)⁢eη⁢T⁢(x)⁢dx.Z^{\prime}(\eta)=\int T(x)\,h(x)\,e^{\eta T(x)}\,\mathrm{d}x.

Dividing by Z⁢(η)=eK⁢(η)Z(\eta)=e^{K(\eta)} changes the integrand to T⁢(x)T(x) times the density f⁢(x;η)f(x;\eta) and thus we get

K′⁢(η)=Z′⁢(η)Z⁢(η)=∫T⁢(x)⁢h⁢(x)⁢eη⁢T⁢(x)−K⁢(η)⁢dx=𝔼η⁢(T⁢(X)).K^{\prime}(\eta)=\frac{Z^{\prime}(\eta)}{Z(\eta)}=\int T(x)\,h(x)\,e^{\eta T(x% )-K(\eta)}\,\mathrm{d}x=\mathbb{E}_{\eta}\bigl{(}T(X)\bigr{)}.

Taking derivatives again, we find

K′′⁢(η)=Z′′⁢(η)Z⁢(η)−(Z′⁢(η)Z⁢(η))2,K^{\prime\prime}(\eta)=\frac{Z^{\prime\prime}(\eta)}{Z(\eta)}-\Bigl{(}\frac{Z^% {\prime}(\eta)}{Z(\eta)}\Bigr{)}^{2},

and using the same argument as above, with T⁢(x)2T(x)^{2} in the integrand, we find Z′′⁢(η)/Z⁢(η)=𝔼η⁢(T⁢(X)2)Z^{\prime\prime}(\eta)/Z(\eta)=\mathbb{E}_{\eta}\bigl{(}T(X)^{2}\bigr{)}. Thus we have K′′⁢(η)=𝔼η⁢(T⁢(X)2)−𝔼η⁢(T⁢(X))2=Varη(T⁢(X))K^{\prime\prime}(\eta)=\mathbb{E}_{\eta}(T(X)^{2})-\mathbb{E}_{\eta}(T(X))^{2}% =\mathop{\mathrm{Var}}\nolimits_{\eta}(T(X)) and the proof is complete. ∎

The theorem allows us to compute moments by differentiation: for the Poisson distribution K⁢(η)=eηK(\eta)=e^{\eta} we find K′⁢(η)=K′′⁢(η)=eη=θK^{\prime}(\eta)=K^{\prime\prime}(\eta)=e^{\eta}=\theta, the mean and variance from table A.1, and in exercise 7.4 we consider the exponential and normal distributions. For TnT_{n}, being a sum of nn i.i.d. copies of T⁢(X)T(X), we can use the theorem to find 𝔼η⁢(Tn)=n⁢K′⁢(η)\mathbb{E}_{\eta}(T_{n})=nK^{\prime}(\eta) and Varη(Tn)=n⁢K′′⁢(η)\mathop{\mathrm{Var}}\nolimits_{\eta}(T_{n})=nK^{\prime\prime}(\eta).

7.3 Several Parameters

The normal distribution where both μ\mu and σ2\sigma^{2} are unknown does not fit definition 7.1, since there are two unknown quantities in the exponent. The extension here is to allow the sum of products in the exponent, one for each parameter.

Definition 7.8.

A parametric model (f⁢(x;θ)|θ∈Θ)\bigl{(}f(x;\theta)\!\mathrel{\big{|}}\!\theta\in\Theta\bigr{)} with Θ⊆ℝk\Theta\subseteq\mathbb{R}^{k} is a kk-parameter exponential family, if the density or the probability weights can be written as

f⁢(x;θ)=h⁢(x)⁢exp⁡(∑j=1kηj⁢(θ)⁢Tj⁢(x)−B⁢(θ))f(x;\theta)=h(x)\exp\Bigl{(}\sum_{j=1}^{k}\eta_{j}(\theta)\,T_{j}(x)-B(\theta)% \Bigr{)}

for all xx and all θ∈Θ\theta\in\Theta, where h≥0h\geq 0 and T1,…,TkT_{1},\dots,T_{k} are functions of xx alone and η1,…,ηk\eta_{1},\dots,\eta_{k} and BB are functions of θ\theta alone. The vector (T1⁢(X),…,Tk⁢(X))\bigl{(}T_{1}(X),\dots,T_{k}(X)\bigr{)} is the natural statistic and (η1⁢(θ),…,ηk⁢(θ))\bigl{(}\eta_{1}(\theta),\dots,\eta_{k}(\theta)\bigr{)} is the natural parameter of the family.

The support of the distribution is again {x∣h⁢(x)>0}\{x\mid h(x)>0\} and the proof of lemma 7.5 applies without change: an i.i.d. sample is again a kk-parameter exponential family, where the natural statistic is the vector of the kk sums ∑iTj⁢(Xi)\sum_{i}T_{j}(X_{i}). The cumulant function can be extended to this situation, but we do not need this result here.

Example 7.9.

By expanding the square in the exponent of the density of the normal distribution N⁢(μ,σ2)N(\mu,\sigma^{2}) from table 1.2, we can write

f⁢(x;μ,σ2)=12⁢π⁢σ2⁢exp⁡(−x22⁢σ2+μσ2⁢x−μ22⁢σ2).f(x;\mu,\sigma^{2})=\frac{1}{\sqrt{2\pi\sigma^{2}}}\exp\Bigl{(}-\frac{x^{2}}{2% \sigma^{2}}+\frac{\mu}{\sigma^{2}}\,x-\frac{\mu^{2}}{2\sigma^{2}}\Bigr{)}.

If only μ\mu is unknown, the factor (2⁢π⁢σ2)−1/2(2\pi\sigma^{2})^{-1/2} and the term −x2/(2⁢σ2)-x^{2}/(2\sigma^{2}) contain no unknown quantities and we can define the function h⁢(x)=(2⁢π⁢σ2)−1/2⁢e−x2/(2⁢σ2)h(x)=(2\pi\sigma^{2})^{-1/2}e^{-x^{2}/(2\sigma^{2})}: we have found a one-parameter exponential family with natural statistic T⁢(x)=xT(x)=x, natural parameter η⁢(μ)=μ/σ2\eta(\mu)=\mu/\sigma^{2} and B⁢(μ)=μ2/(2⁢σ2)B(\mu)=\mu^{2}/(2\sigma^{2}). If both parameters are unknown, we have to interpret the term −x2/(2⁢σ2)-x^{2}/(2\sigma^{2}) as a second product and we get a two-parameter exponential family with h⁢(x)=1h(x)=1, natural statistic (x,x2)(x,x^{2}), natural parameter (μ/σ2,−1/(2⁢σ2))\bigl{(}\mu/\sigma^{2},\,-1/(2\sigma^{2})\bigr{)} and B⁢(θ)=μ2/(2⁢σ2)+12⁢log⁡(2⁢π⁢σ2)B(\theta)=\mu^{2}/(2\sigma^{2})+\frac{1}{2}\log(2\pi\sigma^{2}). For a sample, the natural statistic (∑ixi,∑ixi2)\bigl{(}\sum_{i}x_{i},\sum_{i}x_{i}^{2}\bigr{)} is equivalent to the pair (x¯,∑i(xi−x¯)2)\bigl{(}\bar{x},\sum_{i}(x_{i}-\bar{x})^{2}\bigr{)} used to derive the maximum likelihood estimator in example 5.4.

The gamma distribution with both parameters unknown is a second example of a two-parameter exponential family. If we write xα−1=e(α−1)⁢log⁡xx^{\alpha-1}=e^{(\alpha-1)\log x}, then the natural statistic is (log⁡x,x)(\log x,\,x) and the details can be worked out using exercise 7.3. For a sample of size nn, the natural statistic is the pair (∑ilog⁡xi,∑ixi)\bigl{(}\sum_{i}\log x_{i},\sum_{i}x_{i}\bigr{)}, which is different from the pair of moments used in example 4.9. This is the first sign that the moment estimator is missing some information. (We will see in lecture 8 that this is indeed the case.)

For the completeness statements in lecture 10 it is important that the natural parameter varies in as many directions as there are natural statistics. Mathematically, this can be expressed as the following condition.

Definition 7.10.

A one-parameter exponential family is of full rank, if the set η⁢(Θ)={η⁢(θ)∣θ∈Θ}\eta(\Theta)=\{\eta(\theta)\mid\theta\in\Theta\} of natural parameter values contains a non-degenerate open interval. A kk-parameter exponential family is of full rank, if the set {(η1⁢(θ),…,ηk⁢(θ))|θ∈Θ}\bigl{\{}\bigl{(}\eta_{1}(\theta),\dots,\eta_{k}(\theta)\bigr{)}\mathrel{\big{% |}}\theta\in\Theta\bigr{\}} contains a non-empty open subset of ℝk\mathbb{R}^{k}.

In the one-parameter case, the condition is almost always satisfied, since the image of an interval Θ\Theta under a continuous, non-constant η\eta is an interval with more than one point; all the one-parameter examples above are of full rank. For several parameters, the condition can be more interesting: the natural parameter values of the normal distribution fill the open half-plane {(η1,η2)∣η2<0}\{(\eta_{1},\eta_{2})\mid\eta_{2}<0\} (example 7.9) and the natural parameter values of the gamma distribution fill an open quadrant (exercise 7.3), but the natural parameter values of N⁢(θ,θ2)N(\theta,\theta^{2}) with θ>0\theta>0, where the standard deviation equals the mean, form a curve in ℝ2\mathbb{R}^{2} which does not contain an open set and thus the two-parameter family in this case is not of full rank (exercise 7.7).

7.4 What the Structure Buys

The definitions in this lecture have applications in many areas. In lecture 8 we see that the natural statistic TnT_{n} of a sample is sufficient and lemma 7.5 is essentially the whole proof. In lecture 10 we see that full rank implies that TnT_{n} is complete and we use the Lehmann–Scheffe theorem to find the best unbiased estimators. In lecture 11 we find that the Fisher information for the natural parameter equals K′′⁢(η)K^{\prime\prime}(\eta). In lecture 13 we see that the estimators which achieve the Cramer–Rao bound are, up to affine transformations, natural statistics. Finally, in lecture 26 we find that conjugate priors arise because the likelihood (7.2) depends on the data only via TnT_{n} and nn.

Exponential families also satisfy the regularity conditions from lectures 11, 16 and 17: the support of the distribution is fixed, the density is smooth in θ\theta whenever η\eta and BB are, and we have differentiation under the integral sign as in the proof of theorem 7.7. The log-likelihood ℓ⁢(η)=∑ilog⁡h⁢(xi)+η⁢Tn−n⁢K⁢(η)\ell(\eta)=\sum_{i}\log h(x_{i})+\eta T_{n}-nK(\eta) has second derivative −n⁢K′′⁢(η)=−n⁢Varη(T⁢(X))<0-nK^{\prime\prime}(\eta)=-n\mathop{\mathrm{Var}}\nolimits_{\eta}\bigl{(}T(X)% \bigr{)}<0, since T⁢(X)T(X) is not constant, and thus is strictly concave in η\eta. For this reason, the likelihood equation from exercise 7.5 has at most one root and this root is then the maximum likelihood estimator. This shows that the uniqueness hypothesis of the consistency theorem from lecture 16 holds for the whole class.

Summary.
  • •
    ​

    A one-parameter exponential family has density of the form f⁢(x;θ)=h⁢(x)⁢exp⁡(η⁢(θ)⁢T⁢(x)−B⁢(θ))f(x;\theta)=h(x)\exp\bigl{(}\eta(\theta)T(x)-B(\theta)\bigr{)}, where the natural statistic is TT, the natural parameter is η\eta and the support of the distribution does not depend on θ\theta. Examples of exponential families include the Bernoulli, Poisson, exponential, normal and gamma distributions. The uniform distribution is not an exponential family.

  • •
    ​

    An i.i.d. sample from an exponential family is again an exponential family. In this case, the natural statistic is Tn=∑iT⁢(Xi)T_{n}=\sum_{i}T(X_{i}) and the likelihood depends on the data only through TnT_{n}.

  • •
    ​

    In the natural parametrisation f⁢(x;η)=h⁢(x)⁢exp⁡(η⁢T⁢(x)−K⁢(η))f(x;\eta)=h(x)\exp(\eta T(x)-K(\eta)) the cumulant function KK can be used to find the moments of the natural statistic: we have 𝔼η⁢(T⁢(X))=K′⁢(η)\mathbb{E}_{\eta}(T(X))=K^{\prime}(\eta) and Varη(T⁢(X))=K′′⁢(η)\mathop{\mathrm{Var}}\nolimits_{\eta}(T(X))=K^{\prime\prime}(\eta).

  • •
    ​

    A family is of full rank, if the values of the natural parameter corresponding to an open set of the correct dimension can be achieved. This condition is required for completeness in lecture 10.

  • •
    ​

    By construction, exponential families satisfy the regularity conditions of the later theory and the log-likelihood function is strictly concave in the natural parameter.

Exercise 7.1.

Let XX be geometrically distributed with parameter p∈(0,1)p\in(0,1), i.e. f⁢(x;p)=(1−p)x−1⁢pf(x;p)=(1-p)^{x-1}p for x∈{1,2,…}x\in\{1,2,\dots\}, as in exercises 4.2 and 5.1.

  1. 1.
    ​

    Show that this distribution is a one-parameter exponential family and identify the functions hh, TT, η\eta and BB. What is the natural statistic of an i.i.d. sample X1,…,XnX_{1},\dots,X_{n}?

  2. 2.
    ​

    Determine the natural parameter space 𝒩︀\mathcal{N} and the cumulant function KK. Check that B⁢(p)=K⁢(η⁢(p))B(p)=K\bigl{(}\eta(p)\bigr{)}. Is this family of full rank?

  3. 3.
    ​

    Using theorem 7.7, determine the mean and variance of XX and compare the result with table A.1.

Exercise 7.2.

Let XX be binomially distributed with a known number mm of trials and unknown success probability θ∈(0,1)\theta\in(0,1), i.e. f⁢(x;θ)=(mx)⁢θx⁢(1−θ)m−xf(x;\theta)=\binom{m}{x}\theta^{x}(1-\theta)^{m-x} for x∈{0,1,…,m}x\in\{0,1,\dots,m\}.

  1. 1.
    ​

    Show that this distribution is a one-parameter exponential family, with the same natural parameter as the Bernoulli distribution from example 7.2, and identify the functions hh, TT and BB.

  2. 2.
    ​

    Determine the natural parameter space and the cumulant function.

  3. 3.
    ​

    Using theorem 7.7, determine the mean and variance of XX, and the mean and variance of the natural statistic TnT_{n} for an i.i.d. sample of size nn.

Exercise 7.3.

Let XX be gamma distributed, X∼Gamma⁢(α,β)X\sim\text{Gamma}(\alpha,\beta) with shape α>0\alpha>0 and rate β>0\beta>0, both unknown, i.e. with density given by f⁢(x;α,β)=βα⁢xα−1⁢e−β⁢x/Γ⁢(α)f(x;\alpha,\beta)=\beta^{\alpha}x^{\alpha-1}e^{-\beta x}/\Gamma(\alpha) for x>0x>0, as given in table A.2.

  1. 1.
    ​

    Show that this model is a two-parameter exponential family, as given in definition 7.8, with h⁢(x)=1h(x)=1 for x>0x>0 and h⁢(x)=0h(x)=0 otherwise, natural statistic (log⁡x,x)(\log x,\,x), natural parameter (α−1,−β)(\alpha-1,\,-\beta), and function B⁢(α,β)=log⁡Γ⁢(α)−α⁢log⁡βB(\alpha,\beta)=\log\Gamma(\alpha)-\alpha\log\beta.

  2. 2.
    ​

    Show that the set of natural parameter values is the open quadrant

    {(η1,η2)∣η1>−1,η2<0}\{(\eta_{1},\eta_{2})\mid\eta_{1}>-1,\ \eta_{2}<0\}

    and thus the family has full rank as given in definition 7.10.

  3. 3.
    ​

    Assume that the shape α\alpha is known, as in exercise 5.2. Show that then the factor xα−1x^{\alpha-1} can be incorporated into the function hh to yield a one-parameter exponential family with natural statistic T⁢(x)=xT(x)=x and natural parameter −β-\beta.

Exercise 7.4.
  1. 1.
    ​

    For the exponential distribution with rate θ\theta, use the cumulant function K⁢(η)=−log⁡(−η)K(\eta)=-\log(-\eta) from the text to find the mean and variance of an observation.

  2. 2.
    ​

    For the normal distribution with known variance σ2\sigma^{2} and unknown mean μ\mu, where hh, TT and η\eta are as in example 7.9, use the method of completing the square to show that K⁢(η)=σ2⁢η2/2K(\eta)=\sigma^{2}\eta^{2}/2 with 𝒩︀=ℝ\mathcal{N}=\mathbb{R} and use this to derive the mean and variance of an observation. Check that your result satisfies B⁢(μ)=K⁢(η⁢(μ))B(\mu)=K\bigl{(}\eta(\mu)\bigr{)}.

Exercise 7.5.

Let X1,…,XnX_{1},\dots,X_{n} be an i.i.d. sample from a one-parameter exponential family in its natural parametrisation, with 𝒩︀\mathcal{N} an open interval and Varη(T⁢(X1))>0\mathop{\mathrm{Var}}\nolimits_{\eta}(T(X_{1}))>0 for all η∈𝒩︀\eta\in\mathcal{N}.

  1. 1.
    ​

    Show that the likelihood equation ℓ′⁢(η)=0\ell^{\prime}(\eta)=0 from lecture 5 can be written as

    1n⁢∑i=1nT⁢(xi)=𝔼η⁢(T⁢(X1)),\frac{1}{n}\sum_{i=1}^{n}T(x_{i})=\mathbb{E}_{\eta}\bigl{(}T(X_{1})\bigr{)},

    i.e. the maximum likelihood estimator is the parameter value under which the population mean of the natural statistic equals its sample mean.

  2. 2.
    ​

    Show that every root of the likelihood equation is the unique maximum likelihood estimator for η\eta.

  3. 3.
    ​

    Apply the equation from the first part to the Poisson distribution and to the Bernoulli distribution, and use theorem 5.6 to recover the estimators found in exercise 5.7 and example 5.3.

  4. 4.
    ​

    Explain how the first part of this exercise accounts for the coincidences between the method of moments and maximum likelihood which we noted in lecture 4, and why no such coincidence is to be expected for the gamma distribution of exercise 7.3.

Exercise 7.6.

Let XX be a random variable with density

f⁢(x;η)=h⁢(x)⁢exp⁡(η⁢T⁢(x)−K⁢(η))f(x;\eta)=h(x)\exp\bigl{(}\eta T(x)-K(\eta)\bigr{)}

given by definition 7.6, and let ss be such that η+s∈𝒩︀\eta+s\in\mathcal{N}.

  1. 1.
    ​

    Show that 𝔼η⁢(es⁢T⁢(X))=exp⁡(K⁢(η+s)−K⁢(η))\mathbb{E}_{\eta}\bigl{(}e^{sT(X)}\bigr{)}=\exp\bigl{(}K(\eta+s)-K(\eta)\bigr{)}, i.e. that M⁢(s)=K⁢(η+s)−K⁢(η)M(s)=K(\eta+s)-K(\eta) is the logarithm of the moment generating function of T⁢(X)T(X).

  2. 2.
    ​

    Let η\eta be an interior point of 𝒩︀\mathcal{N}. Then, since MM is differentiable at s=0s=0, we can find M′⁢(0)M^{\prime}(0) and M′′⁢(0)M^{\prime\prime}(0) from the first part of this exercise. Use these results to give an alternative proof of theorem 7.7.

  3. 3.
    ​

    Using the first part of the question, show that for the Poisson distribution with parameter θ\theta we have 𝔼θ⁢(zX)=eθ⁢(z−1)\mathbb{E}_{\theta}\bigl{(}z^{X}\bigr{)}=e^{\theta(z-1)} for all z>0z>0. This is the same identity as in exercise 5.7.

Exercise 7.7.

Consider the normal distribution N⁢(θ,θ2)N(\theta,\theta^{2}) with mean and standard deviation both equal to θ>0\theta>0.

  1. 1.
    ​

    Using example 7.9, show that this model is a two-parameter exponential family with natural statistic and natural parameter given by

    T⁢(x)=(x,x2)andη⁢(θ)=(1/θ,−1/(2⁢θ2)).T(x)=(x,x^{2})\qquad\text{and}\qquad\eta(\theta)=\bigl{(}1/\theta,\,-1/(2% \theta^{2})\bigr{)}.
  2. 2.
    ​

    Show that as θ\theta varies, the natural parameter values follow the curve η2=−η12/2\eta_{2}=-\eta_{1}^{2}/2 with η1>0\eta_{1}>0. Conclude that this family is not of full rank in the sense of definition 7.10.