2.1 Maximum-Entropy Models

Suppose we want to construct a probability model that constrains the expectations of all and only JJ functions of the random variable:

𝔼pˇ⁢[t1⁢(𝑿ˇ)]=μ1,\displaystyle\mathbb{E}_{\check{p}}{\mathopen{}\mathclose{{}\left[t_{1}({\bm{% \check{X}}})}\right]}=\mu_{1},
𝔼pˇ⁢[t2⁢(𝑿ˇ)]=μ2,\displaystyle\mathbb{E}_{\check{p}}{\mathopen{}\mathclose{{}\left[t_{2}({\bm{% \check{X}}})}\right]}=\mu_{2},
…,\displaystyle\ldots,
𝔼pˇ⁢[tJ⁢(𝑿ˇ)]=μj.\displaystyle\mathbb{E}_{\check{p}}{\mathopen{}\mathclose{{}\left[t_{J}({\bm{% \check{X}}})}\right]}=\mu_{j}.

For reasons that will become clear, we call these functions t1,t2,…,tJt_{1},t_{2},\ldots,t_{J} the sufficient statistics††margin: sufficient statistics for the distribution. Conceptually, we are not (yet) claiming to know μ1,μ2,…,μj\mu_{1},\mu_{2},\ldots,\mu_{j}; only that our model constrains the expected values of the sufficient statistics and nothing else. In other words, we seek the distribution that is consistent with these constraints but otherwise makes the no assumptions about the data; that is, the distribution with maximum entropy††margin: maximum-entropy distribution , subject to the constraints.

To solve this constrained optimization problem, we write out the Lagrangian, including the constraint (α\alpha) that the distribution integrate to unity:

ℒ︀⁢(pˇ)=Hpˇ⁢(𝑿ˇ)+α⁢(∫𝒙ˇpˇ⁢(𝒙ˇ;𝜼)⁢d𝒙ˇ−1)+∑j=1Jηj⁢(𝔼pˇ⁢[tj⁢(𝑿ˇ)]−μj)=+∫𝒙ˇpˇ⁢(𝒙ˇ;𝜼)⁢(−log⁡pˇ⁢(𝒙ˇ;𝜼)+α+∑j=1Jηj⁢tj⁢(𝒙ˇ))⁢d𝒙ˇ.\begin{split}\mathcal{L}(\check{p})&{}=H_{\check{p}}({\bm{\check{X}}})+\alpha% \mathopen{}\mathclose{{}\left(\int_{\bm{\check{x}}{}}{\check{p}\mathopen{}% \mathclose{{}\left(\bm{\check{x}};\bm{\eta}}\right)}\mathop{}\!\mathrm{d}{\bm{% \check{x}}{}}-1}\right)+\sum_{j=1}^{J}\eta_{j}\bigg{(}\mathbb{E}_{\check{p}}{% \mathopen{}\mathclose{{}\left[t_{j}({\bm{\check{X}}})}\right]}-\mu_{j}\bigg{)}% \\ &{}\stackrel{{\scriptstyle\text{+}}}{{=}}\int_{\bm{\check{x}}{}}{\check{p}% \mathopen{}\mathclose{{}\left(\bm{\check{x}};\bm{\eta}}\right)}\mathopen{}% \mathclose{{}\left(-\log{\check{p}\mathopen{}\mathclose{{}\left(\bm{\check{x}}% ;\bm{\eta}}\right)}+\alpha+\sum_{j=1}^{J}\eta_{j}t_{j}(\bm{\check{x}})}\right)% \mathop{}\!\mathrm{d}{\bm{\check{x}}{}}.\end{split}

This functional is minimized for functions pˇ⁢(𝒙;𝜼){\check{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{};\bm{\eta}}\right)} that obey the Euler-Lagrange equation (see Chapter A), which in this case simply prescribes that the derivative of the integrand with respect to pˇ\check{p} must be zero. If we define t0(𝒙) . . =1t_{0}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}})\mathrel{\vbox{\hbox{.}\hbox{.} }}=1 and η0 . . =α−1\eta_{0}\mathrel{\vbox{\hbox{.}\hbox{.} }}=\alpha-1, we can write the constraints all together as M(𝒙) . . =1+∑j=0Jηjtj(𝒙)M({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{x}})\mathrel{\vbox{\hbox{.}\hbox{.} }}=1+\sum_{j=0}^{J}\eta_{j}t_{j}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}) (the extra 1 will come in handy). Suppressing 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}-dependence wherever possible for brevity:

0=set∂∂pˇ⁢pˇ⁢(−log⁡pˇ+M)=(−log⁡pˇ+M)+pˇ⁢(−1pˇ)=−log⁡pˇ+(M−1)⟹pˇ⁢(𝒙;𝜼)=exp⁡{∑j=0Jηj⁢tj⁢(𝒙)}.\begin{split}0&{}\stackrel{{\scriptstyle\text{set}}}{{=}}\frac{\partial{}}{% \partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\check{p}}}}\check{p}\mathopen{}\mathclose{{}\left(-\log\check{p}+M}% \right)=\mathopen{}\mathclose{{}\left(-\log\check{p}+M}\right)+\check{p}% \mathopen{}\mathclose{{}\left(-\frac{1}{\check{p}}}\right)=-\log\check{p}+(M-1% )\\ \implies{\check{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{};\bm{\eta}}\right)% }&{}=\exp\mathopen{}\mathclose{{}\left\{\sum_{j=0}^{J}\eta_{j}t_{j}({\color[% rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})}% \right\}.\end{split}

Since solving for η0\eta_{0} requires enforcing the normalization constraint, it is perhaps more common to see the maximum-entropy distribution written as:

equation (2.1) (2.1)
pˇ⁢(𝒙;𝜼)=1Z⁢(𝜼)⁢exp⁡{∑j=1Jηj⁢tj⁢(𝒙)}=1Z⁢(𝜼)⁢exp⁡{𝜼⋅𝒕⁢(𝒙)}{\check{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{};\bm{\eta}}\right)}=\frac{1}{Z(% \bm{\eta})}\exp\mathopen{}\mathclose{{}\left\{\sum_{j=1}^{J}\eta_{j}t_{j}({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x% }})}\right\}=\frac{1}{Z(\bm{\eta})}\exp\mathopen{}\mathclose{{}\left\{\bm{\eta% }\cdot\bm{t}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}})}\right\}

with the partition function Z⁢(𝜼)Z(\bm{\eta}) defined so as the make the RHS integrate or sum to unity.

2.1.1 A second relationship between parameterizations

Eq. 2.1 is a perfectly valid and complete description of a maximum-entropy distribution. Indeed, we refer to it as the natural parameterization, and the Lagrange multipliers η1,η2,…,ηJ\eta_{1},\eta_{2},\ldots,\eta_{J} as the natural parameters††margin: natural parameters . Nevertheless, it is equivalent and sometimes more useful to express the distribution in terms of the constraint values, μ1,μ2,…,μJ\mu_{1},\mu_{2},\ldots,\mu_{J}—what we call the moment parameters††margin: moment parameters —rather than the natural parameters. It seems that to do so we must compute expectations of the sufficient statistics, t1,t2,…⁢tJt_{1},t_{2},\ldots t_{J}.

It turns out that we do not. In fact, conveniently, rather than computing integrals (expectations), we can compute derivatives. This is a consequence of a special property of the logarithm of the partition function, Z⁢(𝜼)Z(\bm{\eta}):

g(𝜼) . . =dd⁢𝜼logZ(𝜼)=1Z⁢(𝜼)⁢dd⁢𝜼⁢∫xˇexp⁡{𝜼⋅𝒕⁢(xˇ)}⁢dxˇ=∫xˇ𝒕⁢(xˇ)⁢1Z⁢(𝜼)⁢exp⁡{𝜼⋅𝒕⁢(xˇ)}⁢dxˇ=𝔼pˇ⁢[𝒕⁢(𝑿ˇ)]=𝝁\begin{split}g(\bm{\eta})\mathrel{\vbox{\hbox{.}\hbox{.} }}=\frac{\mathrm{d}{}}{\mathrm{d}{\bm{\eta}}}\log Z(\bm{\eta})&{}=\frac{1}{Z(% \bm{\eta})}\frac{\mathrm{d}{}}{\mathrm{d}{\bm{\eta}}}\int_{\check{x}{}}\exp% \mathopen{}\mathclose{{}\left\{\bm{\eta}\cdot\bm{t}(\check{x})}\right\}\mathop% {}\!\mathrm{d}{\check{x}{}}\\ &{}=\int_{\check{x}{}}\bm{t}(\check{x})\frac{1}{Z(\bm{\eta})}\exp\mathopen{}% \mathclose{{}\left\{\bm{\eta}\cdot\bm{t}(\check{x})}\right\}\mathop{}\!\mathrm% {d}{\check{x}{}}\\ &{}=\mathbb{E}_{\check{p}}{\mathopen{}\mathclose{{}\left[\bm{t}({\bm{\check{X}% }})}\right]}=\bm{\mu}\end{split}

This shows that the moment parameters can be mapped into the natural parameters by the derivative of the log-partition function. To show that the reverse is also true, i.e. that gg is invertible, we show that its derivative is positive definite (the multivariate equivalent of monotonically increasing):

dd⁢𝜼T⁢g⁢(𝜼)=dd⁢𝜼T⁢∫xˇ𝒕⁢(xˇ)⁢1Z⁢(𝜼)⁢exp⁡{𝜼⋅𝒕⁢(xˇ)}⁢dxˇ=∫xˇ𝒕⁢(xˇ)⁢(1Z⁢(𝜼)⁢𝒕⁢(xˇ)T⁢exp⁡{𝜼⋅𝒕⁢(xˇ)}−1Z2⁢(𝜼)⁢∂Z∂𝜼T⁢(𝜼)⁢exp⁡{𝜼⋅𝒕⁢(xˇ)})⁢dxˇ=∫xˇ𝒕⁢(xˇ)⁢1Z⁢(𝜼)⁢(𝒕⁢(xˇ)T−1Z⁢(𝜼)⁢Z⁢dd⁢𝜼T⁢log⁡Z⁢(𝜼))⁢exp⁡{𝜼⋅𝒕⁢(xˇ)}⁢dxˇ=∫xˇ𝒕⁢(xˇ)⁢1Z⁢(𝜼)⁢(𝒕⁢(xˇ)−𝝁)T⁢exp⁡{𝜼⋅𝒕⁢(xˇ)}⁢dxˇ=𝔼pˇ⁢[𝒕⁢(𝑿ˇ)⁢𝒕⁢(𝑿ˇ)T]−𝝁⁢𝝁T=Covpˇ⁢[𝒕⁢(𝑿ˇ)].\begin{split}\frac{\mathrm{d}{}}{\mathrm{d}{\bm{\eta}}^{\text{T}}}g(\bm{\eta})% &{}=\frac{\mathrm{d}{}}{\mathrm{d}{\bm{\eta}}^{\text{T}}}\int_{\check{x}{}}\bm% {t}(\check{x})\frac{1}{Z(\bm{\eta})}\exp\mathopen{}\mathclose{{}\left\{\bm{% \eta}\cdot\bm{t}(\check{x})}\right\}\mathop{}\!\mathrm{d}{\check{x}{}}\\ &{}=\int_{\check{x}{}}\bm{t}(\check{x})\mathopen{}\mathclose{{}\left(\frac{1}{% Z(\bm{\eta})}\bm{t}(\check{x})^{\text{T}}\exp\mathopen{}\mathclose{{}\left\{% \bm{\eta}\cdot\bm{t}(\check{x})}\right\}-\frac{1}{Z^{2}(\bm{\eta})}\frac{% \partial{Z}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\eta}}}^{\text{T}}}(\bm{\eta})\exp\mathopen% {}\mathclose{{}\left\{\bm{\eta}\cdot\bm{t}(\check{x})}\right\}}\right)\mathop{% }\!\mathrm{d}{\check{x}{}}\\ &{}=\int_{\check{x}{}}\bm{t}(\check{x})\frac{1}{Z(\bm{\eta})}\mathopen{}% \mathclose{{}\left(\bm{t}(\check{x})^{\text{T}}-\frac{1}{Z(\bm{\eta})}Z\frac{% \mathrm{d}{}}{\mathrm{d}{\bm{\eta}}^{\text{T}}}\log Z(\bm{\eta})}\right)\exp% \mathopen{}\mathclose{{}\left\{\bm{\eta}\cdot\bm{t}(\check{x})}\right\}\mathop% {}\!\mathrm{d}{\check{x}{}}\\ &{}=\int_{\check{x}{}}\bm{t}(\check{x})\frac{1}{Z(\bm{\eta})}\mathopen{}% \mathclose{{}\left(\bm{t}(\check{x})-\bm{\mu}}\right)^{\text{T}}\exp\mathopen{% }\mathclose{{}\left\{\bm{\eta}\cdot\bm{t}(\check{x})}\right\}\mathop{}\!% \mathrm{d}{\check{x}{}}\\ &{}=\mathbb{E}_{\check{p}}{\mathopen{}\mathclose{{}\left[\bm{t}({\bm{\check{X}% }})\bm{t}({\bm{\check{X}}})^{\text{T}}}\right]}-\bm{\mu}\bm{\mu}^{\text{T}}\\ &{}=\text{Cov}_{\check{p}}{\mathopen{}\mathclose{{}\left[\bm{t}({\bm{\check{X}% }})}\right]}.\end{split}

Covariance matrices are positive definite, so g⁢(𝜼)g(\bm{\eta}) is invertible, and the log-partition function itself is convex.††margin: What about positive semi-definite covariances?

The log-partition function is useful enough in its own right to deserve a symbol, for which we use A(𝜼) . . =logZ(𝜼)A(\bm{\eta})\mathrel{\vbox{\hbox{.}\hbox{.} }}=\log Z(\bm{\eta}). In summary, we have

equation (2.2) (2.2)
𝝁=g⁢(𝜼)=∂A∂𝜼⁢(𝜼),\displaystyle\bm{\mu}=g(\bm{\eta})=\frac{\partial{A}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\eta}}}}(\bm{% \eta}),
𝜼=g−1⁢(𝝁).\displaystyle\bm{\eta}=g^{-1}(\bm{\mu}).

Example: the normal distribution.

Consider the distribution of data on the real line with fixed mean μ1\mu_{1} and variance σ\sigma; or equivalently, mean μ1\mu_{1} and expected square μ2\mu_{2}. That is, our two constraints are t1⁢(x)=xt_{1}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}x})={\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}x} and t2⁢(x)=x2t_{2}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}x})={\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}x}^{2}. From Eq. 2.1, we have

pˇ⁢(x;𝜼)=1Z⁢(𝜼)⁢exp⁡{η1⁢x+η2⁢x2}.{\check{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}x};\bm{\eta}}\right)}=\frac{1}{Z(\bm{% \eta})}\exp\mathopen{}\mathclose{{}\left\{\eta_{1}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}x}+\eta_{2}{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}x}^{2}}\right\}.

Notice that η2<0\eta_{2}<0, otherwise probability would grow without bound with x{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}x}{}.

In fact, this equation already tells us that Xˇ{\check{X}} is normally distributed, since the Gaussian is the unique distribution on the real line with a second-order log probability. But we will pretend not to know this. To impose the constraint that the distribution integrate to unity, we complete the square (first line) and apply the formula for a Gaussian integral††margin: Exercise LABEL:ex: :

Z⁢(𝜼)=set∫−∞∞exp⁡{η2⁢xˇ2+η1⁢xˇ}⁢dxˇ=∫−∞∞exp⁡{η2⁢(xˇ+η12⁢η2)2−η124⁢η2}⁢dxˇ=−τ2⁢η2⁢exp⁡{−η124⁢η2}⟹A⁢(𝜼)=log⁡Z⁢(𝜼)=12⁢log⁡τ−12⁢log⁡(−2⁢η2)−η124⁢η2.\begin{split}Z(\bm{\eta})&{}\stackrel{{\scriptstyle\text{set}}}{{=}}\int_{-% \infty}^{\infty}\exp\mathopen{}\mathclose{{}\left\{\eta_{2}\check{x}^{2}+\eta_% {1}\check{x}}\right\}\mathop{}\!\mathrm{d}{\check{x}{}}=\int_{-\infty}^{\infty% }\exp\mathopen{}\mathclose{{}\left\{\eta_{2}\mathopen{}\mathclose{{}\left(% \check{x}+\frac{\eta_{1}}{2\eta_{2}}}\right)^{2}-\frac{\eta_{1}^{2}}{4\eta_{2}% }}\right\}\mathop{}\!\mathrm{d}{\check{x}{}}\\ &{}=\sqrt{\frac{-\tau}{2\eta_{2}}}\exp\mathopen{}\mathclose{{}\left\{-\frac{% \eta_{1}^{2}}{4\eta_{2}}}\right\}\\ \implies A(\bm{\eta})&{}=\log Z(\bm{\eta})=\frac{1}{2}\log\tau-\frac{1}{2}\log% (-2\eta_{2})-\frac{\eta_{1}^{2}}{4\eta_{2}}.\end{split}

Now we can proceed in either of two ways. The simplest is to differentiate the log-partition function:

μ1\displaystyle\mu_{1}
=set∂A∂η1⁢(𝜼)=−η12⁢η2\displaystyle{}\stackrel{{\scriptstyle\text{set}}}{{=}}\frac{\partial{A}}{% \partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\eta_{1}}}}(\bm{\eta})=-\frac{\eta_{1}}{2\eta_{2}}
μ2\displaystyle\mu_{2}
=set∂A∂η2⁢(𝜼)=−12⁢η2+η124⁢η22=η12−2⁢η24⁢η22\displaystyle{}\stackrel{{\scriptstyle\text{set}}}{{=}}\frac{\partial{A}}{% \partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\eta_{2}}}}(\bm{\eta})=-\frac{1}{2\eta_{2}}+\frac{\eta_{1}^{2}}{4% \eta_{2}^{2}}=\frac{\eta_{1}^{2}-2\eta_{2}}{4\eta_{2}^{2}}

Alternatively, we can apply expectations. The constraint on the first moment implies that

μ1=set1Z⁢(𝜼)⁢∫−∞∞xˇ⁢exp⁡{η2⁢xˇ2+η1⁢xˇ}⁢dxˇ=−2⁢η2τ⁢∫−∞∞xˇ⁢exp⁡{η2⁢(xˇ+η12⁢η2)2}⁢dxˇ=−2⁢η2τ⁢∫−∞∞(w⁢exp⁡{η2⁢w2}−η12⁢η2⁢exp⁡{η2⁢w2})⁢dw=−η12⁢η2,\begin{split}\mu_{1}&{}\stackrel{{\scriptstyle\text{set}}}{{=}}\frac{1}{Z(\bm{% \eta})}\int_{-\infty}^{\infty}\check{x}\exp\mathopen{}\mathclose{{}\left\{\eta% _{2}\check{x}^{2}+\eta_{1}\check{x}}\right\}\mathop{}\!\mathrm{d}{\check{x}{}}% =\sqrt{\frac{-2\eta_{2}}{\tau}}\int_{-\infty}^{\infty}\check{x}\exp\mathopen{}% \mathclose{{}\left\{\eta_{2}\mathopen{}\mathclose{{}\left(\check{x}+\frac{\eta% _{1}}{2\eta_{2}}}\right)^{2}}\right\}\mathop{}\!\mathrm{d}{\check{x}{}}\\ &{}=\sqrt{\frac{-2\eta_{2}}{\tau}}\int_{-\infty}^{\infty}\mathopen{}\mathclose% {{}\left(w\exp\mathopen{}\mathclose{{}\left\{\eta_{2}w^{2}}\right\}-\frac{\eta% _{1}}{2\eta_{2}}\exp\mathopen{}\mathclose{{}\left\{\eta_{2}w^{2}}\right\}}% \right)\mathop{}\!\mathrm{d}{w{}}\\ &{}=-\frac{\eta_{1}}{2\eta_{2}},\end{split}

where the first integral vanishes by odd symmetry of w⁢exp⁡{η2⁢w2}w\exp\mathopen{}\mathclose{{}\left\{\eta_{2}w^{2}}\right\} (recalling again that η2<0\eta_{2}<0), and the second is another Gaussian integral. This is the same relationship obtained from the derivative of the log-partition function. Application of the constraint on the second moment is left as an exercise.††margin: Exercise LABEL:ex:

To express the distribution in the moment parameterization, we invert the relationship between 𝝁\bm{\mu} and 𝜼\bm{\eta}:

η1\displaystyle\eta_{1}
=μ1μ2−μ12,\displaystyle{}=\frac{\mu_{1}}{\mu_{2}-\mu_{1}^{2}},
η2\displaystyle\eta_{2}
=−12⁢(μ2−μ12).\displaystyle{}=-\frac{1}{2(\mu_{2}-\mu_{1}^{2})}.

Inserting these into the natural parameterization yields

pˇ⁢(x;𝝁)=1−τ⁢(μ2−μ12)⁢exp⁡{2⁢μ12⁢(μ2−μ12)⁢x−12⁢(μ2−μ12)⁢x2−μ122⁢(μ2−μ12)}.{\check{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}x};\bm{\mu}}\right)}=\sqrt{\frac{1}{-% \tau(\mu_{2}-\mu_{1}^{2})}}\exp\mathopen{}\mathclose{{}\left\{\frac{2\mu_{1}}{% 2(\mu_{2}-\mu_{1}^{2})}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}x}-\frac{1}{2(\mu_{2}-\mu_{1}^{2})}{\color[rgb]% {.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}x}^{2}-\frac{\mu% _{1}^{2}}{2(\mu_{2}-\mu_{1}^{2})}}\right\}.

Conversion into the standard form of the Gaussian is left as an exercise for the reader.††margin: Exercise LABEL:ex:

Example: the categorical distribution.

Consider the distribution of outcomes of the toss of a K{K}-sided die, with fixed mean/probability of each side πj\pi_{j}. We represent an outcome (realization) with a one-hot vector 𝒙ˇ\bm{\check{x}}, and the collection of means with the vector 𝝅\bm{\pi}. Note that this amounts to K−1{K}-1, rather than K{K}, constraints, because the mean vector 𝝅\bm{\pi} must sum to 1: one of the constraints is determined by the others. Still, for mathematical convenience, let us introduce an extra parameter but fix its value to zero: ηK . . =0\eta_{{K}}\mathrel{\vbox{\hbox{.}\hbox{.} }}=0. Then by Eq. 2.1, the maximum-entropy distribution is

pˇ⁢(𝒙;𝜼)=1Z⁢(𝜼)⁢exp⁡{∑k=1K−1ηk⁢xk}=1Z⁢(𝜼)⁢exp⁡{∑k=1Kηk⁢xk}.{\check{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{};\bm{\eta}}\right)}=\frac{1}{Z(% \bm{\eta})}\exp\mathopen{}\mathclose{{}\left\{\sum_{k=1}^{{K}-1}\eta_{k}{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}x}_{k% }}\right\}=\frac{1}{Z(\bm{\eta})}\exp\mathopen{}\mathclose{{}\left\{\sum_{k=1}% ^{{K}}\eta_{k}{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}x}_{k}}\right\}.

The normalizer requires

Z⁢(𝜼)=set∑𝒙ˇexp⁡{∑k=1Kηk⋅xˇk}=∑k=1Kexp⁡{ηk}.Z(\bm{\eta})\stackrel{{\scriptstyle\text{set}}}{{=}}\sum_{\bm{\check{x}}{}}% \exp\mathopen{}\mathclose{{}\left\{\sum_{k=1}^{{K}}\eta_{k}\cdot\check{x}_{k}}% \right\}=\sum_{k=1}^{{K}}\exp\mathopen{}\mathclose{{}\left\{\eta_{k}}\right\}.

We again consider both methods for computing the mean constraint. The derivative of the log-partition function is

𝝅=set∂A∂𝜼⁢(𝜼)=exp⁡{𝜼}∑k=1Kexp⁡{ηk}=softmax{𝜼}\bm{\pi}\stackrel{{\scriptstyle\text{set}}}{{=}}\frac{\partial{A}}{\partial{{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \eta}}}}(\bm{\eta})=\frac{\exp\mathopen{}\mathclose{{}\left\{\bm{\eta}}\right% \}}{\sum_{k=1}^{{K}}\exp\mathopen{}\mathclose{{}\left\{\eta_{k}}\right\}}=% \operatorname*{softmax}\mathopen{}\mathclose{{}\left\{\bm{\eta}}\right\}

Alternatively, we compute the expectation:

𝝅=set𝔼pˇ⁢[𝑿ˇ]=∑𝒙ˇ𝒙ˇ⁢1Z⁢(𝜼)⁢exp⁡{𝜼⋅𝒙ˇ}=exp⁡{𝜼}∑k=1Kexp⁡{ηk}.\bm{\pi}\stackrel{{\scriptstyle\text{set}}}{{=}}\mathbb{E}_{\check{p}}{% \mathopen{}\mathclose{{}\left[{\bm{\check{X}}}}\right]}=\sum_{\bm{\check{x}}{}% }\bm{\check{x}}\frac{1}{Z(\bm{\eta})}\exp\mathopen{}\mathclose{{}\left\{\bm{% \eta}\cdot\bm{\check{x}}}\right\}=\frac{\exp\mathopen{}\mathclose{{}\left\{\bm% {\eta}}\right\}}{\sum_{k=1}^{{K}}\exp\mathopen{}\mathclose{{}\left\{\eta_{k}}% \right\}}.

The problem of inverting the softmax might appear ill-posed, but recall that the last natural parameter ηK=0\eta_{{K}}=0 is fictitious. In particular,

πK=exp⁡{ηK}∑k=1Kexp⁡{ηk}=1∑k=1Kexp⁡{ηk}⟹𝝅=exp⁡{𝜼}⁢πK⟹𝜼=log⁡(𝝅πK).\pi_{{K}}=\frac{\exp\mathopen{}\mathclose{{}\left\{\eta_{{K}}}\right\}}{\sum_{% k=1}^{{K}}\exp\mathopen{}\mathclose{{}\left\{\eta_{k}}\right\}}=\frac{1}{\sum_% {k=1}^{{K}}\exp\mathopen{}\mathclose{{}\left\{\eta_{k}}\right\}}\implies\bm{% \pi}=\exp\mathopen{}\mathclose{{}\left\{\bm{\eta}}\right\}\pi_{{K}}\implies\bm% {\eta}=\log\mathopen{}\mathclose{{}\left(\frac{\bm{\pi}}{\pi_{{K}}}}\right).

With this relationship we can reparameterize our maximum-entropy distribution:

pˇ⁢(𝒙;𝝅)=πK⁢exp⁡{∑k=1Klog⁡(πkπK)⁢xk}=∏k=1Kπkxk.{\check{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{};\bm{\pi}}\right)}=\pi_{{K}}% \exp\mathopen{}\mathclose{{}\left\{\sum_{k=1}^{{K}}\log\mathopen{}\mathclose{{% }\left(\frac{\pi_{k}}{\pi_{{K}}}}\right){\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}x}_{k}}\right\}=\prod_{k=1}^{{K}}\pi_{k}% ^{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}x}% _{k}}.

This is a standard form for the categorical distribution.

family 𝒕⁢(𝒙ˇ)\bm{t}(\bm{\check{x}}) A⁢(𝜼)A(\bm{\eta}) g⁢(𝜼)g(\bm{\eta}) g−1⁢(𝝁)g^{-1}(\bm{\mu}) h⁢(𝒙ˇ)h(\bm{\check{x}})
univariate Gaussian [xˇxˇ2]\begin{bmatrix}\check{x}\\ \check{x}^{2}\end{bmatrix} 12⁢log⁡τ−2⁢η2−η124⁢η2\frac{1}{2}\log\frac{\tau}{-2\eta_{2}}-\frac{\eta_{1}^{2}}{4\eta_{2}} [−η12⁢η2η12−2⁢η24⁢η22]\begin{bmatrix}-\frac{\eta_{1}}{2\eta_{2}}\\ \frac{\eta_{1}^{2}-2\eta_{2}}{4\eta_{2}^{2}}\end{bmatrix} [μ1μ2−μ12−1/2μ2−μ12]\begin{bmatrix}\frac{\mu_{1}}{\mu_{2}-\mu_{1}^{2}}\\ \frac{-1/2}{\mu_{2}-\mu_{1}^{2}}\end{bmatrix} 1
categorical 𝒙ˇ\bm{\check{x}} log⁡(∑k=1Kexp⁡{ηk})\log\mathopen{}\mathclose{{}\left(\sum_{k=1}^{{K}}\exp\mathopen{}\mathclose{{}% \left\{\eta_{k}}\right\}}\right) softmax{𝜼}\operatorname*{softmax}\mathopen{}\mathclose{{}\left\{\bm{\eta}}\right\} log⁡(𝝅πK)\log\mathopen{}\mathclose{{}\left(\frac{\bm{\pi}}{\pi_{{K}}}}\right) 1
exponential xˇ\check{x} −log⁡(−η)-\log\mathopen{}\mathclose{{}\left(-\eta}\right) −1η-\frac{1}{\eta} −1μ-\frac{1}{\mu} 1

Example: the exponential distribution.

Consider the distribution of waiting times on the interval [0,∞)[0,\infty), where events occur at a mean rate of μ\mu. So t⁢(x)=xt({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}x}% )={\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}x}. Then Eq. 2.1 tells us that

pˇ⁢(x;η)=1Z⁢(η)⁢exp⁡{η⁢x}.{\check{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}x};\eta}\right)}=\frac{1}{Z(\eta)}\exp% \mathopen{}\mathclose{{}\left\{\eta{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}x}}\right\}.

As usual, we could stop here, since it is well known that the unique distribution defined on the non-negative reals with a first-order log probability is the exponential distribution. We can derive this by applying the constraints. The normalizer requires

Z⁢(η)=set∫0∞exp⁡{η⁢xˇ}⁢dxˇ=−1η.Z(\eta)\stackrel{{\scriptstyle\text{set}}}{{=}}\int_{0}^{\infty}\exp\mathopen{% }\mathclose{{}\left\{\eta\check{x}}\right\}\mathop{}\!\mathrm{d}{\check{x}{}}=% -\frac{1}{\eta}.

Since Z>0Z>0, η<0\eta<0. Likewise, the constraint on the mean requires

μ=set𝔼pˇ⁢[Xˇ]=1Z⁢(η)⁢∫0∞xˇ⁢exp⁡{η⁢xˇ}⁢dxˇ=1Z⁢(η)(1ηxˇexp{ηxˇ}|0∞−∫0∞1ηexp{ηxˇ}dxˇ)=1Z⁢(η)⁢η2,\begin{split}\mu\stackrel{{\scriptstyle\text{set}}}{{=}}\mathbb{E}_{\check{p}}% {\mathopen{}\mathclose{{}\left[{\check{X}}}\right]}&{}=\frac{1}{Z(\eta)}\int_{% 0}^{\infty}\check{x}\exp\mathopen{}\mathclose{{}\left\{\eta\check{x}}\right\}% \mathop{}\!\mathrm{d}{\check{x}{}}\\ &{}=\frac{1}{Z(\eta)}\mathopen{}\mathclose{{}\left(\mathopen{}\mathclose{{}% \left.\frac{1}{\eta}\check{x}\exp\mathopen{}\mathclose{{}\left\{\eta\check{x}}% \right\}}\right\rvert_{0}^{\infty}-\int_{0}^{\infty}\frac{1}{\eta}\exp% \mathopen{}\mathclose{{}\left\{\eta\check{x}}\right\}\mathop{}\!\mathrm{d}{% \check{x}{}}}\right)=\frac{1}{Z(\eta)\eta^{2}},\end{split}

where the final equality follows because η<0\eta<0. Hence μ=−1/η\mu=-1/\eta. Alternatively, differentiating the log-partition function,

μ=set∂A∂η⁢(η)=−1η,\mu\stackrel{{\scriptstyle\text{set}}}{{=}}\frac{\partial{A}}{\partial{{\color% [rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\eta}}}(% \eta)=-\frac{1}{\eta},

yields the same result. The moment parameterization of this distribution is therefore

pˇ⁢(x;μ)=1μ⁢exp⁡{−xμ}.{\check{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}x};\mu}\right)}=\frac{1}{\mu}\exp% \mathopen{}\mathclose{{}\left\{-\frac{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}x}}{\mu}}\right\}.

which is indeed the exponential distribution.

Example: the Poisson distribution.

The distribution of counts for exponentially distributed waiting times is Poisson, so we are licensed to conclude that the maximum-entropy distribution for counts is Poisson. But it would be nice to arrive at this conclusion without such foreknowledge.††margin: FINISH THIS CHAPTER