Lecture 7 Exponential Families
In lectures 4 and 5 we repeatedly collapsed products of densities to expressions which summarised the data using only one or two sums. This lecture explains the reason for this coincidence: almost all families of models considered here have the same algebraic shape, and models of this shape are called exponential families. This shape allows us to define the data summary used in the likelihood, the moments of this summary, and the smoothness required for the later theory, and many of the results in the following weeks are almost automatic for this class.
7.1 Recognising the Shape
We start by considering a familiar distribution and then rewrite its probability weights so that the data and the parameter are as far apart as possible. For the Bernoulli distribution with we can write
Here we have written the weights as the exponential of the product of one function of and one function of , minus a function of only . The Poisson weights and the exponential density can be written in a similar way: we have and , up to a factor depending on . We will give a name to this structure.
A parametric model with is a one-parameter exponential family, if the density or the probability weights can be written as
for all and all , where and are functions of alone and and are functions of alone. In particular, the support is and does not depend on . The statistic is called the natural statistic and is called the natural parameter of the family.
Since the exponential is always positive, is positive if and only if is. This proves the statement about the support. The representation is not unique, since we can replace by and by for any constant . Finally, since the density must integrate (or sum) to one, the function is determined by the other three functions: we have
These rewritings give the following examples.
The Bernoulli distribution with is a one-parameter exponential family with for , natural statistic , natural parameter , the log-odds of success, and .
The Poisson distribution with is a one-parameter exponential family with for , natural statistic , natural parameter and .
The exponential distribution with rate is a one-parameter exponential family with for and otherwise, natural statistic , natural parameter and .
The normal distribution is also an exponential family, but since both parameters are unknown, we need to use the two-parameter version of the definition from section 7.3.
The uniform distribution on is not an exponential family, since the support of the distribution changes with the parameter. In contrast, the support of an exponential family is always fixed. This is the same reason why the likelihood equation was useless in example 5.5. Similarly, the location model from section 5.3 is not an exponential family, despite the fact that its support is all of ; we do not prove this.
The shape survives the passage from one observation to a sample.
Since the observations are independent, the joint density is the product of the marginal densities. Combining the exponents using the rule gives the required form. This completes the proof. ∎
Read as a function of , the joint density is the likelihood, and thus the log-likelihood of an exponential family sample is
The first term does not involve and thus, except for the term involving , the likelihood depends on the data only via the single number . This is an extension of example 4.3 to the whole class; in lecture 8 we will see in which sense the value contains all the information about .
7.2 The Natural Parametrisation and the Cumulant Function
Since the parameter enters the density only via , it is often convenient to use as the parameter. In this case, the normalising function (7.1) could be given a name, since its derivatives are the moments of the natural statistic.
Let and be as in definition 7.1. Then the natural parametrisation of the family is given by
where the cumulant function (also called the log-partition function) is given by
where the sum replaces the integral for discrete models, and the natural parameter space is the set of all such that the integral is finite.
Since makes integrate to one, the formula defines a probability distribution for every , whether or not is of the form ; comparing with (7.1) shows that and , so that the original family sits inside this larger one. For the three examples of section 7.1 we find with for the Bernoulli distribution, with for the Poisson distribution and with for the exponential distribution. In all three cases is an open interval and is a bijection onto , making the natural parametrisation a relabelling in the sense of theorem 5.6.
The first two derivatives of give the mean and the variance of . The name of is derived from the fact that is the cumulant generating function of , as given in exercise 7.6.
Let be an interior point of the natural parameter space . Then, in the natural parametrisation,
Set , so that . Since is not identically zero and , we have and thus the division in the proof is justified. Since is in the interior of , we can take derivatives inside the integral as many times as required; this is a result from analysis which we use without proof. Taking derivatives once, we get
Dividing by changes the integrand to times the density and thus we get
Taking derivatives again, we find
and using the same argument as above, with in the integrand, we find . Thus we have and the proof is complete. ∎
7.3 Several Parameters
The normal distribution where both and are unknown does not fit definition 7.1, since there are two unknown quantities in the exponent. The extension here is to allow the sum of products in the exponent, one for each parameter.
A parametric model with is a -parameter exponential family, if the density or the probability weights can be written as
for all and all , where and are functions of alone and and are functions of alone. The vector is the natural statistic and is the natural parameter of the family.
The support of the distribution is again and the proof of lemma 7.5 applies without change: an i.i.d. sample is again a -parameter exponential family, where the natural statistic is the vector of the sums . The cumulant function can be extended to this situation, but we do not need this result here.
By expanding the square in the exponent of the density of the normal distribution from table 1.2, we can write
If only is unknown, the factor and the term contain no unknown quantities and we can define the function : we have found a one-parameter exponential family with natural statistic , natural parameter and . If both parameters are unknown, we have to interpret the term as a second product and we get a two-parameter exponential family with , natural statistic , natural parameter and . For a sample, the natural statistic is equivalent to the pair used to derive the maximum likelihood estimator in example 5.4.
The gamma distribution with both parameters unknown is a second example of a two-parameter exponential family. If we write , then the natural statistic is and the details can be worked out using exercise 7.3. For a sample of size , the natural statistic is the pair , which is different from the pair of moments used in example 4.9. This is the first sign that the moment estimator is missing some information. (We will see in lecture 8 that this is indeed the case.)
For the completeness statements in lecture 10 it is important that the natural parameter varies in as many directions as there are natural statistics. Mathematically, this can be expressed as the following condition.
A one-parameter exponential family is of full rank, if the set of natural parameter values contains a non-degenerate open interval. A -parameter exponential family is of full rank, if the set contains a non-empty open subset of .
In the one-parameter case, the condition is almost always satisfied, since the image of an interval under a continuous, non-constant is an interval with more than one point; all the one-parameter examples above are of full rank. For several parameters, the condition can be more interesting: the natural parameter values of the normal distribution fill the open half-plane (example 7.9) and the natural parameter values of the gamma distribution fill an open quadrant (exercise 7.3), but the natural parameter values of with , where the standard deviation equals the mean, form a curve in which does not contain an open set and thus the two-parameter family in this case is not of full rank (exercise 7.7).
7.4 What the Structure Buys
The definitions in this lecture have applications in many areas. In lecture 8 we see that the natural statistic of a sample is sufficient and lemma 7.5 is essentially the whole proof. In lecture 10 we see that full rank implies that is complete and we use the Lehmann–Scheffe theorem to find the best unbiased estimators. In lecture 11 we find that the Fisher information for the natural parameter equals . In lecture 13 we see that the estimators which achieve the Cramer–Rao bound are, up to affine transformations, natural statistics. Finally, in lecture 26 we find that conjugate priors arise because the likelihood (7.2) depends on the data only via and .
Exponential families also satisfy the regularity conditions from lectures 11, 16 and 17: the support of the distribution is fixed, the density is smooth in whenever and are, and we have differentiation under the integral sign as in the proof of theorem 7.7. The log-likelihood has second derivative , since is not constant, and thus is strictly concave in . For this reason, the likelihood equation from exercise 7.5 has at most one root and this root is then the maximum likelihood estimator. This shows that the uniqueness hypothesis of the consistency theorem from lecture 16 holds for the whole class.
-
•
A one-parameter exponential family has density of the form , where the natural statistic is , the natural parameter is and the support of the distribution does not depend on . Examples of exponential families include the Bernoulli, Poisson, exponential, normal and gamma distributions. The uniform distribution is not an exponential family.
-
•
An i.i.d. sample from an exponential family is again an exponential family. In this case, the natural statistic is and the likelihood depends on the data only through .
-
•
In the natural parametrisation the cumulant function can be used to find the moments of the natural statistic: we have and .
-
•
A family is of full rank, if the values of the natural parameter corresponding to an open set of the correct dimension can be achieved. This condition is required for completeness in lecture 10.
-
•
By construction, exponential families satisfy the regularity conditions of the later theory and the log-likelihood function is strictly concave in the natural parameter.
Let be geometrically distributed with parameter , i.e. for , as in exercises 4.2 and 5.1.
-
1.
Show that this distribution is a one-parameter exponential family and identify the functions , , and . What is the natural statistic of an i.i.d. sample ?
-
2.
Determine the natural parameter space and the cumulant function . Check that . Is this family of full rank?
- 3.
Let be binomially distributed with a known number of trials and unknown success probability , i.e. for .
-
1.
Show that this distribution is a one-parameter exponential family, with the same natural parameter as the Bernoulli distribution from example 7.2, and identify the functions , and .
-
2.
Determine the natural parameter space and the cumulant function.
-
3.
Using theorem 7.7, determine the mean and variance of , and the mean and variance of the natural statistic for an i.i.d. sample of size .
Let be gamma distributed, with shape and rate , both unknown, i.e. with density given by for , as given in table A.2.
-
1.
Show that this model is a two-parameter exponential family, as given in definition 7.8, with for and otherwise, natural statistic , natural parameter , and function .
-
2.
Show that the set of natural parameter values is the open quadrant
and thus the family has full rank as given in definition 7.10.
-
3.
Assume that the shape is known, as in exercise 5.2. Show that then the factor can be incorporated into the function to yield a one-parameter exponential family with natural statistic and natural parameter .
-
1.
For the exponential distribution with rate , use the cumulant function from the text to find the mean and variance of an observation.
-
2.
For the normal distribution with known variance and unknown mean , where , and are as in example 7.9, use the method of completing the square to show that with and use this to derive the mean and variance of an observation. Check that your result satisfies .
Let be an i.i.d. sample from a one-parameter exponential family in its natural parametrisation, with an open interval and for all .
-
1.
Show that the likelihood equation from lecture 5 can be written as
i.e. the maximum likelihood estimator is the parameter value under which the population mean of the natural statistic equals its sample mean.
-
2.
Show that every root of the likelihood equation is the unique maximum likelihood estimator for .
- 3.
- 4.
Let be a random variable with density
given by definition 7.6, and let be such that .
-
1.
Show that , i.e. that is the logarithm of the moment generating function of .
-
2.
Let be an interior point of . Then, since is differentiable at , we can find and from the first part of this exercise. Use these results to give an alternative proof of theorem 7.7.
-
3.
Using the first part of the question, show that for the Poisson distribution with parameter we have for all . This is the same identity as in exercise 5.7.
Consider the normal distribution with mean and standard deviation both equal to .