3.3 Exponential families

So far we have seen how certain judicious choices of conditional distributions lead to tractable marginalization and “posteriorization” (inversion with Bayes’s rule). Here we approach from something like the opposite direction, allowing any exponential family, but then considering, somewhat informally, under what conditions marginalization and “posteriorization” are tractable.

3.3.1 Posteriorization (Bayes inversion)

Suppose we have an emission function in the exponential family,

equation (3.62) (3.62)
logp^(𝒚|𝒙;𝜽)=+𝜼(𝒙)⋅𝑻(𝒚)−A(𝜼(𝒙))=[𝜼⁢(𝒙)−A⁢(𝜼⁢(𝒙))]⋅[𝑻⁢(𝒚)1]\log{\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{}\middle|{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{};\bm{\theta}}% \right)}\stackrel{{\scriptstyle\text{+}}}{{=}}\bm{\eta}({\color[rgb]{.75,0,.25% }\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})\cdot\bm{T}({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y% }})-A(\bm{\eta}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb% }{.75,0,.25}\bm{x}}))=\begin{bmatrix}\bm{\eta}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})\\ -A(\bm{\eta}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}))\end{bmatrix}\cdot\begin{bmatrix}\bm{T}({\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}})\\ 1\end{bmatrix}

(the omitted constant may depend on 𝒚{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% y}}), and we wish to construct a prior that is conjugate to the conditioning variable, 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}. Then we need to choose the sufficient statistics 𝑼⁢(𝒙)\bm{U}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}) for the prior (with natural parameters 𝝉\bm{\tau}) so as to “mimic the form of the likelihood”; i.e., we need the 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}-dependent vector in the likelihood to match the sufficient stastistics of the prior, up to an affine transformation:

equation (3.63) (3.63)
log⁡p^⁢(𝒙;𝜽)=+𝝉⋅𝑼⁢(𝒙)+log⁡h⁢(𝒙)⟹[𝜼⁢(𝒙)−A⁢(𝜼⁢(𝒙))]=+𝐌⁢𝑼⁢(𝒙)\log{\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{};\bm{\theta}}\right)}\stackrel{% {\scriptstyle\text{+}}}{{=}}\bm{\tau}\cdot\bm{U}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})+\log h({\color[rgb% ]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})% \implies\begin{bmatrix}\bm{\eta}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})\\ -A(\bm{\eta}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}))\end{bmatrix}\stackrel{{\scriptstyle\text{+}}}{{=}}\mathbf{% M}\bm{U}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}})

The matrix 𝐌\mathbf{M} allows some flexibility; in particular, it allows the emission’s log-partition function, A⁢(𝜼⁢(𝒙))A(\bm{\eta}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}})), as well as its natural parameters, 𝜼⁢(𝒙)\bm{\eta}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}), to be constructed from the same set of prior sufficient statistics, 𝑼⁢(𝒙)\bm{U}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}). Since the dimensions of the sufficient statistics for the prior and emission need not be related to each other, 𝐌\mathbf{M} will not in general be square. We only require equality up to an additive constant to account for any 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}-independent terms in the natural parameters or log-partition function. Note finally that the base measure, h⁢(𝒙)h({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{x}}), will be treated entirely separately; we see why next.

When 𝑼⁢(𝒙)\bm{U}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}) is defined this way, Bayes’s rule (up to the normalizer) amounts to

equation (3.64) (3.64)
logp^(𝒙|𝒚;𝜽)=+𝐌⁢𝑼⁢(𝒙)⋅[𝑻⁢(𝒚)1]+𝝉⋅𝑼⁢(𝒙)+log⁡h⁢(𝒙)=+𝑼⁢(𝒙)⋅(𝝉+𝐌T⁢[𝑻⁢(𝒚)1])+log⁡h⁢(𝒙).\begin{split}\log{\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{}\middle|{\color[% rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{};% \bm{\theta}}\right)}&{}\stackrel{{\scriptstyle\text{+}}}{{=}}\mathbf{M}\bm{U}(% {\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}})\cdot\begin{bmatrix}\bm{T}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}})\\ 1\end{bmatrix}+\bm{\tau}\cdot\bm{U}({\color[rgb]{.75,0,.25}\definecolor[named]% {pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})+\log h({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})\\ &{}\stackrel{{\scriptstyle\text{+}}}{{=}}\bm{U}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})\cdot\mathopen{}% \mathclose{{}\left(\bm{\tau}+\mathbf{M}^{\text{T}}\begin{bmatrix}\bm{T}({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y% }})\\ 1\end{bmatrix}}\right)+\log h({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}).\end{split}

Clearly, the sufficient statistics (𝑼⁢(𝒙)\bm{U}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}})) and base measure (h⁢(𝒙)h({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{x}})) for the prior will play the same roles in the posterior. That leaves the parenthetical term to be the natural parameters of the posterior, which we will call 𝜻⁢(𝒚)\bm{\zeta}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{y}}):

equation (3.65) (3.65)
𝜻(𝒚) . . =𝝉+𝐌T[𝑻⁢(𝒚)1].\bm{\zeta}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{y}})\mathrel{\vbox{\hbox{.}\hbox{.} }}=\bm{\tau}+\mathbf{M}^{\text{T}}\begin{bmatrix}\bm{T}({\color[rgb]{.75,0,.25% }\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}})\\ 1\end{bmatrix}.

Hence, incorporating new observations (moving from prior to posterior) amounts merely to updating the natural parameters—simply by adding a term derived from the emission’s sufficient statistics. In fine, Bayes inversion in exponential families amounts to a linear operation in natural parameters (of the prior) and sufficient statistics (of the emission).

Multivariate normal emissions.

For example, for multivariate normal emissions 𝒀^{\bm{\hat{Y}}} with means affine in 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}, the natural parameters and sufficient statistics are

equation (3.66) (3.66)
𝜼⁢(𝒙)=[𝚺y^|x^−1⁢(𝐂⁢𝒙+𝒄)−12⁢vec⁢[𝚺y^|x^−1]],\displaystyle\bm{\eta}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})=\begin{bmatrix}\mathbf{{\Sigma}}^{-1}_% {\hat{y}|\hat{x}}\mathopen{}\mathclose{{}\left(\mathbf{{C}}{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}+\bm{c}}% \right)\\ -\frac{1}{2}\text{vec}\mathopen{}\mathclose{{}\left[\mathbf{{\Sigma}}^{-1}_{% \hat{y}|\hat{x}}}\right]\end{bmatrix},
𝑻⁢(𝒚)=[𝒚vec⁢[𝒚⁢𝒚T]],\displaystyle\bm{T}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}% {rgb}{.75,0,.25}\bm{y}})=\begin{bmatrix}{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}\\ \text{vec}\mathopen{}\mathclose{{}\left[{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}^{\text{T}}}\right]% \end{bmatrix},

and the log-partition function is

A⁢(𝜼⁢(𝒙))=12⁢(𝐂⁢𝒙+𝒄)T⁢𝚺y^|x^−1⁢(𝐂⁢𝒙+𝒄)+12⁢log⁡|𝚺y^|x^|=+12⁢𝒙T⁢𝐂T⁢𝚺y^|x^−1⁢𝐂⁢𝒙+𝒄T⁢𝚺y^|x^−1⁢𝐂⁢𝒙\begin{split}A(\bm{\eta}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}))&{}=\frac{1}{2}\mathopen{}\mathclose{{% }\left(\mathbf{{C}}{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{% rgb}{.75,0,.25}\bm{x}}+\bm{c}}\right)^{\text{T}}\mathbf{{\Sigma}}^{-1}_{\hat{y% }|\hat{x}}\mathopen{}\mathclose{{}\left(\mathbf{{C}}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}+\bm{c}}\right)+% \frac{1}{2}\log\mathopen{}\mathclose{{}\left\lvert\mathbf{{\Sigma}}_{\hat{y}|% \hat{x}}}\right\rvert\\ &{}\stackrel{{\scriptstyle\text{+}}}{{=}}\frac{1}{2}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}^{\text{T}}\mathbf{{% C}}^{\text{T}}\mathbf{{\Sigma}}^{-1}_{\hat{y}|\hat{x}}\mathbf{{C}}{\color[rgb]% {.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}+\bm{c}^{% \text{T}}\mathbf{{\Sigma}}^{-1}_{\hat{y}|\hat{x}}\mathbf{{C}}{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}\end{split}

Therefore, following Eq. 3.63, we can choose1010 10 This choice is sufficient—and sensible—but not necessary, since the sufficient statistics are determined only up to an invertible transformation. the sufficient statistics of the prior and posterior according to

[𝜼⁢(𝒙)−A⁢(𝜼⁢(𝒙))]=+[𝚺y^|x^−1⁢𝐂𝟎𝟎𝟎−𝒄T⁢𝚺y^|x^−1⁢𝐂−12⁢vec⁢[𝐂T⁢𝚺y^|x^−1⁢𝐂]T]⏟𝐌⁢[𝒙vec⁢[𝒙⁢𝒙T]]⏟𝑼⁢(𝒙).\begin{split}\begin{bmatrix}\bm{\eta}({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})\\ -A(\bm{\eta}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}))\end{bmatrix}\stackrel{{\scriptstyle\text{+}}}{{=}}% \underbrace{\begin{bmatrix}\mathbf{{\Sigma}}^{-1}_{\hat{y}|\hat{x}}\mathbf{{C}% }&\mathbf{0}\\ \mathbf{0}&\mathbf{0}\\ -\bm{c}^{\text{T}}\mathbf{{\Sigma}}^{-1}_{\hat{y}|\hat{x}}\mathbf{{C}}&-\frac{% 1}{2}\text{vec}\mathopen{}\mathclose{{}\left[\mathbf{{C}}^{\text{T}}\mathbf{{% \Sigma}}^{-1}_{\hat{y}|\hat{x}}\mathbf{{C}}}\right]^{\text{T}}\end{bmatrix}}_{% \mathbf{M}}\underbrace{\begin{bmatrix}{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}\\ \text{vec}\mathopen{}\mathclose{{}\left[{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}^{\text{T}}}\right]% \end{bmatrix}}_{\bm{U}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})}.\end{split}

Notice that 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}} and 𝒙⁢𝒙T{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{x}}^{\text{T}} are sufficient statistics for a multivariate normal distribution (in 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}), so evidently the conjugate prior and corresponding posterior are, like the emission, multivariate normal distributions.

As noted above, the only update required is to the natural parameters. We recall that the natural parameters of a multivariate Gaussian,

p^⁢(𝒙;𝜽)=𝒩︀⁢(𝝁x^,𝚺x^),{\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{};\bm{\theta}}\right)}=\mathcal{% N}\mathopen{}\mathclose{{}\left(\bm{\mu}_{\hat{x}},\>\mathbf{{\Sigma}}_{\hat{x% }}}\right),

are

𝝉=[𝚺x^−1⁢𝝁x^−12⁢vec⁢[𝚺x^−1]].\bm{\tau}=\begin{bmatrix}\mathbf{{\Sigma}}^{-1}_{\hat{x}}\bm{\mu}_{\hat{x}}\\ -\frac{1}{2}\text{vec}\mathopen{}\mathclose{{}\left[\mathbf{{\Sigma}}^{-1}_{% \hat{x}}}\right]\end{bmatrix}.

Therefore, interpreting Eq. 3.65 (which applies to all exponential families with conjugate priors) in light of the present case (to wit, multivariate normal emissions), we find

𝜻⁢(𝒚)=[𝚺x^−1⁢𝝁x^−12⁢vec⁢[𝚺x^−1]]⏟𝝉+[𝐂T⁢𝚺y^|x^−1𝟎−𝐂T⁢𝚺y^|x^−1⁢𝒄𝟎𝟎−12⁢vec⁢[𝐂T⁢𝚺y^|x^−1⁢𝐂]]⏟𝐌T⁢[𝒚vec⁢[𝒚⁢𝒚T]1]⏟[𝑻⁢(𝒚),1]T=[𝚺x^−1⁢𝝁x^+𝐂T⁢𝚺y^|x^−1⁢(𝒚−𝒄)−12⁢vec⁢[𝚺x^−1+𝐂T⁢𝚺y^|x^−1⁢𝐂]],\begin{split}\bm{\zeta}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}})&{}=\underbrace{\begin{bmatrix}\mathbf{% {\Sigma}}^{-1}_{\hat{x}}\bm{\mu}_{\hat{x}}\\ -\frac{1}{2}\text{vec}\mathopen{}\mathclose{{}\left[\mathbf{{\Sigma}}^{-1}_{% \hat{x}}}\right]\end{bmatrix}}_{\bm{\tau}}+\underbrace{\begin{bmatrix}\mathbf{% {C}}^{\text{T}}\mathbf{{\Sigma}}^{-1}_{\hat{y}|\hat{x}}&\mathbf{0}&-\mathbf{{C% }}^{\text{T}}\mathbf{{\Sigma}}^{-1}_{\hat{y}|\hat{x}}\bm{c}\\ \mathbf{0}&\mathbf{0}&-\frac{1}{2}\text{vec}\mathopen{}\mathclose{{}\left[% \mathbf{{C}}^{\text{T}}\mathbf{{\Sigma}}^{-1}_{\hat{y}|\hat{x}}\mathbf{{C}}}% \right]\end{bmatrix}}_{\mathbf{M}^{\text{T}}}\underbrace{\begin{bmatrix}{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y% }}\\ \text{vec}\mathopen{}\mathclose{{}\left[{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}^{\text{T}}}\right]% \\ 1\end{bmatrix}}_{[\bm{T}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}),1]^{\text{T}}}\\ &{}=\begin{bmatrix}\mathbf{{\Sigma}}^{-1}_{\hat{x}}\bm{\mu}_{\hat{x}}+\mathbf{% {C}}^{\text{T}}\mathbf{{\Sigma}}^{-1}_{\hat{y}|\hat{x}}({\color[rgb]{.75,0,.25% }\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}-\bm{c})\\ -\frac{1}{2}\text{vec}\mathopen{}\mathclose{{}\left[\mathbf{{\Sigma}}^{-1}_{% \hat{x}}+\mathbf{{C}}^{\text{T}}\mathbf{{\Sigma}}^{-1}_{\hat{y}|\hat{x}}% \mathbf{{C}}}\right]\end{bmatrix},\end{split}

which corresponds precisely to the well known Bayes inversion of multivariate Gaussians: evidence adds in proportion to confidence, and confidence adds. These match the equations, Eqs. 3.13 and 3.14, that we derived expressly for the normal distribution in Section 3.1.2, by completing the square.

3.3.2 Marginalization

Let us now consider the other main operation on probabilities, marginalization. Now, there are many convenient ways to describe probability distributions. In carrying out Bayes’ rule in the previous section, we used log probabilities and, ultimately, the natural parameters. Marginalizations will turn out to be more conveniently carried out in terms of the moment parameterization, i.e. the expected sufficient statistics. (Recall that in exponential families, there is a bijection between the natural and moment parameters.) This convenience is a consequence of the fact that we can express the expected sufficient statistics of the marginal distribution in terms of the emission and prior probabilities with the law of total expectation:

𝔼𝒀^[𝑻(𝒀^)]=𝔼𝑿^[𝔼𝒀^|𝑿^[𝑻(𝒀^)|𝑿^]].\mathbb{E}_{{\bm{\hat{Y}}}{}}{\mathopen{}\mathclose{{}\left[\bm{T}({\bm{\hat{Y% }}})}\right]}=\mathbb{E}_{{\bm{\hat{X}}}{}}{\mathopen{}\mathclose{{}\left[% \mathbb{E}_{{\bm{\hat{Y}}}{}|{\bm{\hat{X}}}}{\mathopen{}\mathclose{{}\left[\bm% {T}({\bm{\hat{Y}}})\middle|{\bm{\hat{X}}}{}}\right]}}\right]}.

To proceed further, we need to make more assumptions. For example, we might assume that the expected sufficient statistics of the emission are affine functions of the latent variable,

𝔼𝒀^|𝑿^[𝑻(𝒀^)|𝒙]=set𝐏𝒙+𝒑⟹𝔼𝒀^[𝑻(𝒀^)]=𝐏𝔼𝑿^[𝑿^]+𝒑.\mathbb{E}_{{\bm{\hat{Y}}}{}|{\bm{\hat{X}}}}{\mathopen{}\mathclose{{}\left[\bm% {T}({\bm{\hat{Y}}})\middle|{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{}}\right]}\stackrel{{\scriptstyle\text{% set}}}{{=}}\mathbf{P}{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor% }{rgb}{.75,0,.25}\bm{x}}+\bm{p}\implies\mathbb{E}_{{\bm{\hat{Y}}}{}}{\mathopen% {}\mathclose{{}\left[\bm{T}({\bm{\hat{Y}}})}\right]}=\mathbf{P}\mathbb{E}_{{% \bm{\hat{X}}}}{\mathopen{}\mathclose{{}\left[{\bm{\hat{X}}}}\right]}+\bm{p}.

Less restrictively, let the expected sufficient statistics be a quadratic function of the latent variable. Then using Eq. B.15, we have

𝔼𝒀^|𝑿^[𝑻(𝒀^)|𝒙]=set𝒙T⁢𝐏⁢𝒙+𝒑T⁢𝒙+p⟹𝔼𝒀^⁢[𝑻⁢(𝒀^)]=𝔼𝑿^⁢[𝑿^]T⁢𝐏⁢𝔼𝑿^⁢[𝑿^]+𝒑T⁢𝔼𝑿^⁢[𝑿^]+p+tr⁢[𝐏⁢Cov𝑿^⁢[𝑿^]].\begin{split}\mathbb{E}_{{\bm{\hat{Y}}}{}|{\bm{\hat{X}}}}{\mathopen{}% \mathclose{{}\left[\bm{T}({\bm{\hat{Y}}})\middle|{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{}}\right]}&{}% \stackrel{{\scriptstyle\text{set}}}{{=}}{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}^{\text{T}}\mathbf{P}{\color[rgb]% {.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}+\bm{p}^{% \text{T}}{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}+p\\ \implies\mathbb{E}_{{\bm{\hat{Y}}}{}}{\mathopen{}\mathclose{{}\left[\bm{T}({% \bm{\hat{Y}}})}\right]}&{}=\mathbb{E}_{{\bm{\hat{X}}}}{\mathopen{}\mathclose{{% }\left[{\bm{\hat{X}}}}\right]}^{\text{T}}\mathbf{P}\mathbb{E}_{{\bm{\hat{X}}}}% {\mathopen{}\mathclose{{}\left[{\bm{\hat{X}}}}\right]}+\bm{p}^{\text{T}}% \mathbb{E}_{{\bm{\hat{X}}}}{\mathopen{}\mathclose{{}\left[{\bm{\hat{X}}}}% \right]}+p+\text{tr}\mathopen{}\mathclose{{}\left[\mathbf{P}\text{Cov}_{{\bm{% \hat{X}}}}{\mathopen{}\mathclose{{}\left[{\bm{\hat{X}}}}\right]}}\right].\end{split}

Multivariate normal emissions.

For Gaussian random variables, first- and second-order statistics are sufficient; or, put differently, the sufficient statistics are second-order functions of the observations, 𝒚{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% y}}. Therefore, to restrict the expected sufficient statistics to be second-order in the latent variables, 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}, we need to restrict the conditional mean of 𝒀^{\bm{\hat{Y}}} to be affine in 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}. In particular, the conditional expectation of 𝒀^⁢𝒀^T{\bm{\hat{Y}}}{\bm{\hat{Y}}}^{\text{T}} will turn an affine functions of 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}} into a quadratic function of 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}. The conditional covariance of 𝒀^{\bm{\hat{Y}}}, for its part, must be at most quadratic in 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}. In equations, we have

𝔼𝒀^|𝑿^[𝑻(𝒀^)|𝒙]=𝔼𝒀^|𝑿^[[𝒀^𝒀^⁢𝒀^T]|𝒙]=[𝐂⁢𝒙+𝒄(𝐂⁢𝒙+𝒄)⁢(𝐂⁢𝒙+𝒄)T+𝚺y^|x^⁢(𝒙)].\mathbb{E}_{{\bm{\hat{Y}}}{}|{\bm{\hat{X}}}}{\mathopen{}\mathclose{{}\left[\bm% {T}({\bm{\hat{Y}}})\middle|{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{}}\right]}=\mathbb{E}_{{\bm{\hat{Y}}}{}% |{\bm{\hat{X}}}}{\mathopen{}\mathclose{{}\left[\begin{bmatrix}{\bm{\hat{Y}}}\\ {\bm{\hat{Y}}}{\bm{\hat{Y}}}^{\text{T}}\end{bmatrix}\middle|{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{}}\right]% }=\begin{bmatrix}\mathbf{{C}}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{}+\bm{c}\\ \mathopen{}\mathclose{{}\left(\mathbf{{C}}{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{}+\bm{c}}\right)\mathopen{}% \mathclose{{}\left(\mathbf{{C}}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{}+\bm{c}}\right)^{\text{T}}+\mathbf{{% \Sigma}}_{\hat{y}|\hat{x}}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})\end{bmatrix}.

The expected sufficient statistics for the marginal distribution are therefore

𝔼𝒀^⁢[𝑻⁢(𝒀^)]=[𝐂⁢𝝁x^+𝒄𝐂⁢𝔼𝑿^⁢[𝑿^⁢𝑿^T]⁢𝐂T+2⁢𝒄T⁢𝐂⁢𝝁x^+𝒄T⁢𝒄+𝔼𝑿^⁢[𝚺y^|x^⁢(𝑿^)]]=[𝐂⁢𝝁x^+𝒄(𝐂⁢𝝁x^+𝒄)⁢(𝐂⁢𝝁x^+𝒄)T+𝐂⁢𝚺x^⁢𝐂T+𝔼𝑿^⁢[𝚺y^|x^⁢(𝑿^)]].\begin{split}\mathbb{E}_{{\bm{\hat{Y}}}{}}{\mathopen{}\mathclose{{}\left[\bm{T% }({\bm{\hat{Y}}})}\right]}&{}=\begin{bmatrix}\mathbf{{C}}\bm{\mu}_{\hat{x}}+% \bm{c}\\ \mathbf{{C}}\mathbb{E}_{{\bm{\hat{X}}}}{\mathopen{}\mathclose{{}\left[{\bm{% \hat{X}}}{\bm{\hat{X}}}^{\text{T}}}\right]}\mathbf{{C}}^{\text{T}}+2\bm{c}^{% \text{T}}\mathbf{{C}}\bm{\mu}_{\hat{x}}+\bm{c}^{\text{T}}\bm{c}+\mathbb{E}_{{% \bm{\hat{X}}}}{\mathopen{}\mathclose{{}\left[\mathbf{{\Sigma}}_{\hat{y}|\hat{x% }}({\bm{\hat{X}}})}\right]}\end{bmatrix}\\ &{}=\begin{bmatrix}\mathbf{{C}}\bm{\mu}_{\hat{x}}+\bm{c}\\ \mathopen{}\mathclose{{}\left(\mathbf{{C}}\bm{\mu}_{\hat{x}}+\bm{c}}\right)% \mathopen{}\mathclose{{}\left(\mathbf{{C}}\bm{\mu}_{\hat{x}}+\bm{c}}\right)^{% \text{T}}+\mathbf{{C}}\mathbf{{\Sigma}}_{\hat{x}}\mathbf{{C}}^{\text{T}}+% \mathbb{E}_{{\bm{\hat{X}}}}{\mathopen{}\mathclose{{}\left[\mathbf{{\Sigma}}_{% \hat{y}|\hat{x}}({\bm{\hat{X}}})}\right]}\end{bmatrix}.\end{split}

As long as the conditional covariance, 𝚺y^|x^⁢(𝒙)\mathbf{{\Sigma}}_{\hat{y}|\hat{x}}({\color[rgb]{.75,0,.25}\definecolor[named]% {pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}), is second-order in 𝒙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}, the entire expression involves only the first two cumulants of the prior, 𝝁x^\bm{\mu}_{\hat{x}} and 𝚺x^\mathbf{{\Sigma}}_{\hat{x}}. Therefore, if we further assume the prior to be Gaussian, then we can conclude that the marginal is likewise Gaussian, with expected sufficient statistics

𝝁y^\displaystyle\bm{\mu}_{\hat{y}}
=𝐂⁢𝝁x^+𝒄,\displaystyle{}=\mathbf{{C}}\bm{\mu}_{\hat{x}}+\bm{c},
𝚺y^\displaystyle\mathbf{{\Sigma}}_{\hat{y}}
=𝐂⁢𝚺x^⁢𝐂T+𝔼𝑿^⁢[𝚺y^|x^⁢(𝑿^)].\displaystyle{}=\mathbf{{C}}\mathbf{{\Sigma}}_{\hat{x}}\mathbf{{C}}^{\text{T}}% +\mathbb{E}_{{\bm{\hat{X}}}}{\mathopen{}\mathclose{{}\left[\mathbf{{\Sigma}}_{% \hat{y}|\hat{x}}({\bm{\hat{X}}})}\right]}.

When the covariance matrix 𝚺y^|x^\mathbf{{\Sigma}}_{\hat{y}|\hat{x}} is independent of 𝑿^{\bm{\hat{X}}}, these become the equations we derived in Section 3.1.2 using the laws of total expectation and total covariance, Eqs. 3.16 and 3.17.

Let’s relax one assumption and tighten another (although this may make proofs intractable). Let p(x) be any distribution, not necessarily an exponential family (relaxed assumption). But let’s require that