11.1 A duality between generative and discriminative learning

As we have seen, in generative models, expressive power typically comes at the price of ease of inference (and vice versa). So far we have explored three different strategies for this managing trade off: Severely limit expressive power to conjugate or pseudo-conjugate prior distributions in order to allow for exact inference (Chapter 9); allow arbitrarily expressive generative models, but then approximate inference with a separate recogniton model that is arbitrary, but highly parameterized (Sections 10.1 and 10.2), or a generic homogenizer (Section 10.3), or again correct up to simplifying but erroneous independence assumptions (Section 10.4). Now we introduce a fourth strategy: let the latent variables of the model be related to the observed variables by an invertible (and therefore deterministic) transformation. This makes the model marginal p^⁢(𝒚;𝜽){\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{};\bm{\theta}}\right)} computable in closed-form with the standard rules of calculus for changing variables under an integral. Consequently, the marginal relative entropy, DKL⁡{p⁢(𝒀)∥p^⁢(𝒀;𝜽)}\operatorname*{\text{D}_{\text{KL}}}\mathopen{}\mathclose{{}\left\{{p\mathopen% {}\mathclose{{}\left({\bm{Y}}}\right)}\middle\|{\hat{p}\mathopen{}\mathclose{{% }\left({\bm{Y}};\bm{\theta}}\right)}}\right\}, can be descended directly, rather than indirectly via the joint relative entropy, and EM is not needed.

Below we explore two such models in the usual way, starting with simple linear transformations and then moving to more complicated functions. But let us begin with a more abstract formulation, in order to make contact with the unsupervised, discriminative learning problems of Section 7.2. In that section, we related observations 𝒀^{\bm{\hat{Y}}} to (hypothetical) “latent variables” 𝒁^{\bm{\hat{Z}}} via a deterministic, invertible transformation, Eq. 7.23. Here we shall make a similar specification, although to emphasize that this is a generative model, we write the observations as a function of the latent variables, 𝑿^{\bm{\hat{X}}}, rather than vice versa. Still, to exhibit the relationship with the discriminative model, we denote this function as the inverse of some “recognition” function, 𝒅⁢(𝒚,𝜽)\bm{d}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{y}},\bm{\theta}):

equation (11.1) (11.1)
𝒚^=𝒅−1⁢(𝒙,𝜽).\bm{\hat{y}}=\bm{d}^{-1}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}},\bm{\theta}).

For generative models, it is also necessary to specify a prior distribution over the latent variables. Although the discriminative model makes no such specification, its training objective—maximization of the latent (“output”) entropy—will tend to produce independent outputs. Therefore, we shall assume that the generative model’s latent variables are independent—and, for now, nothing else:

equation (11.2) (11.2)
p^𝑿^⁢(𝒙)=∏k=1Kp^X^k⁢(xk)=∏k=1K∂ψk∂x^k⁢(xk).{\hat{p}_{{\bm{\hat{X}}}}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{}}\right)}=\prod_{k% =1}^{{K}}{\hat{p}_{{\hat{X}}_{k}}\mathopen{}\mathclose{{}\left({\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}x_{k}}}\right)}=% \prod_{k=1}^{{K}}\frac{\partial{\psi_{k}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\hat{x}_{k}}}}\mathopen{}% \mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{% rgb}{.75,0,.25}x_{k}}}\right).

In the second equality, we have simply expressed the probability distributions in terms of the cumulative distribution functions (CDFs), ψk⁢(xk)\psi_{k}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}x_{k}}), or more precisely their derivatives, in anticipation of the results below.

In the discriminative models of Section 7.2, rather than being specified explicitly, the distribution of latent variables was inherited from the data distribution, p⁢(𝒚){p\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{}}\right)}, via the change-of-variables formula. Here something like the reverse obtains: The distribution of the (model) observations p^⁢(𝒚;𝜽){\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{};\bm{\theta}}\right)} is inherited, via the deterministic transformation Eq. 11.1 and the change-of-variables formula, from the distribution of latent variables, Eq. 11.2:

equation (11.3) (11.3)
p^⁢(𝒚;𝜽)=p^𝑿^⁢(𝒅⁢(𝒚,𝜽);𝜽)⁢|∂𝒅∂𝒚T⁢(𝒚,𝜽)|=∏k=1K∂ψk∂x^k⁢(dk⁢(𝒚,𝜽))⁢|∂𝒅∂𝒚T⁢(𝒚,𝜽)|.{\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{};\bm{\theta}}\right)}={\hat{p}_% {{\bm{\hat{X}}}}\mathopen{}\mathclose{{}\left(\bm{d}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{},\bm{\theta});\bm{% \theta}}\right)}\mathopen{}\mathclose{{}\left\lvert\frac{\partial{\bm{d}}}{% \partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{y}}}^{\text{T}}}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}},\bm{\theta})}\right\rvert\\ =\prod_{k=1}^{{K}}\frac{\partial{\psi_{k}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\hat{x}_{k}}}}\mathopen{}% \mathclose{{}\left(d_{k}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}},\bm{\theta})}\right)\mathopen{}% \mathclose{{}\left\lvert\frac{\partial{\bm{d}}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}}^{\text{T% }}}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{y}},\bm{\theta})}\right\rvert.

We have recovered Eq. 7.27, but from a different model [5]. Both models use the same map between observations and “latent” variables, given by Eq. 11.1. But the discriminative model proceeds to pass the outputs of this map, 𝒙ˇ\bm{\check{x}}, through squashing functions 𝝍⁢(𝒙)\bm{\psi}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}), whereas the generative model instead interprets 𝝍⁢(𝒙)\bm{\psi}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}) as the (prior) CDFs of those variables, 𝒙^\bm{\hat{x}}. Thus the “outputs” 𝒛ˇ\bm{\check{z}} of the discriminative model are the latent variables of the generative model passed through their own CDFs. If we define 𝒛^ . . =𝝍(𝒙^)\bm{\hat{z}}\mathrel{\vbox{\hbox{.}\hbox{.} }}=\bm{\psi}(\bm{\hat{x}}) for the generative model, then we can say that the generative 𝒁^{\bm{\hat{Z}}} are distributed independently (because 𝑿^{\bm{\hat{X}}} are independent and 𝝍⁢(⋅)\bm{\psi}(\cdot) acts elementwise) and uniformly (because the CDFs exactly flatten the distribution of 𝑿^{\bm{\hat{X}}}). This is consistent with the discriminative objective: maximizing the entropy of 𝒁ˇ{\bm{\check{Z}}} will likewise tend to distribute them independently and uniformly.

Thus, in the deterministic, invertible setting, maximizing mutual information through a discriminative model is equivalent to density estimation in a generative model. It is also useful to re-express the density-estimation problem in terms of 𝑿^{\bm{\hat{X}}}:

equation (11.4) (11.4)
−ℐ︀⁢(𝒀;𝒁ˇ)=DKL⁡{p⁢(𝒀)∥∏kK∂ψk∂xk⁢(dk⁢(𝒀,𝜽))⁢|∂𝒅∂𝒚T⁢(𝒀,𝜽)|}=DKL⁡{pˇ⁢(𝑿;𝜽)∥p^𝑿^⁢(𝑿)},-\mathcal{I}\mathopen{}\mathclose{{}\left({\bm{Y}};{\bm{\check{Z}}}}\right)=% \operatorname*{\text{D}_{\text{KL}}}\mathopen{}\mathclose{{}\left\{{p\mathopen% {}\mathclose{{}\left({\bm{Y}}}\right)}\middle\|\prod_{k}^{{K}}\frac{\partial{% \psi_{k}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}% {rgb}{.75,0,.25}x_{k}}}}(d_{k}({\bm{Y}},\bm{\theta}))\mathopen{}\mathclose{{}% \left\lvert\frac{\partial{\bm{d}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}}^{\text{T}}}({\bm{Y% }},\bm{\theta})}\right\rvert}\right\}=\operatorname*{\text{D}_{\text{KL}}}% \mathopen{}\mathclose{{}\left\{{\check{p}\mathopen{}\mathclose{{}\left({\bm{X}% };\bm{\theta}}\right)}\middle\|{\hat{p}_{{\bm{\hat{X}}}}\mathopen{}\mathclose{% {}\left({\bm{X}}}\right)}}\right\},

where pˇ⁢(𝒙;𝜽){\check{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{};\bm{\theta}}\right)} is the distribution induced by the recognition function applied to the observed data, 𝒅⁢(𝒀,𝜽)\bm{d}({\bm{Y}},\bm{\theta}). The first equality is Eq. 7.26, and the second can be derived simply by noting that relative entropy is invariant under reparameterization (in this case with 𝒅⁢(𝒚,𝜽)\bm{d}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{y}},\bm{\theta})). In light of the second equality, and dropping the language of discriminative and generative models, we can summarize this approach to model fitting like this:

We require a reparameterization of the data, 𝐗ˇ=𝐝⁢(𝐘,𝛉){\bm{\check{X}}}=\bm{d}({\bm{Y}},\bm{\theta}), to be distributed close (in the KL sense) to some factorial prior distribution, ∏k=1K∂ψk∂x^k⁢(xk)\prod_{k=1}^{{K}}\frac{\partial{\psi_{k}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\hat{x}_{k}}}}\mathopen{}% \mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{% rgb}{.75,0,.25}x_{k}}}\right); or, equivalently, to be maximally entropic when “squashed” by ψk⁢(xk)\psi_{k}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}x_{k}}) (for all kk).

11.1.1 InfoMax ICA, revisited

Historically, this equivalence was originally noted [5] for a specific model, InfoMax ICA [3], which we first encountered in Section 7.2. Consider the very simply “generative model” in which the observations are related to the “latent” variables by a square, full-rank matrix:

𝒚^=𝐂⁢𝒙^=𝐃−1⁢𝒙^.\bm{\hat{y}}=\mathbf{{C}}\bm{\hat{x}}=\mathbf{{D}}^{-1}\bm{\hat{x}}.

Substituting this relationship (cf. Eq. 11.1) into Eq. 11.3, we see that the marginal distribution of the observed variables is

equation (11.5) (11.5)
p^⁢(𝒚;𝜽)=∏k=1K∂ψk∂x^k⁢(𝒅kT⁢𝒚)⁢|𝐃|,{\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{};\bm{\theta}}\right)}=\prod_{k=% 1}^{{K}}\frac{\partial{\psi_{k}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor% [named]{pgfstrokecolor}{rgb}{.75,0,.25}\hat{x}_{k}}}}\mathopen{}\mathclose{{}% \left(\bm{d}^{\text{T}}_{k}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}}\right)\mathopen{}\mathclose{{}\left% \lvert\mathbf{{D}}}\right\rvert,

where again ψk\psi_{k} is the CDF of the corresponding “latent” variable, X^k{\hat{X}}_{k}, and 𝒅kT\bm{d}^{\text{T}}_{k} is a row of 𝐃\mathbf{{D}}. Clearly, fitting this marginal density follows the same gradient as in InfoMax ICA, Eq. 7.29.

dd⁢𝐃⁢DKL⁡{p⁢(𝒀)∥∏k=1K∂ψk∂xˇk⁢(𝒅kT⁢𝒀)⁢|𝐃|}=−𝔼𝒀⁢[∑k=1Kdd⁢𝐃⁢log⁡∂ψk∂xˇk⁢(𝒅kT⁢𝒀)]−𝐃−T.\frac{\mathrm{d}{}}{\mathrm{d}{\mathbf{{D}}}}\operatorname*{\text{D}_{\text{KL% }}}\mathopen{}\mathclose{{}\left\{{p\mathopen{}\mathclose{{}\left({\bm{Y}}}% \right)}\middle\|\prod_{k=1}^{{K}}\frac{\partial{\psi_{k}}}{\partial{{\color[% rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\check{x}_{k% }}}}(\bm{d}_{k}^{\text{T}}{\bm{Y}})\mathopen{}\mathclose{{}\left\lvert\mathbf{% {D}}}\right\rvert}\right\}\\ =-\mathbb{E}_{{\bm{Y}}{}}{\mathopen{}\mathclose{{}\left[\sum_{k=1}^{{K}}\frac{% \mathrm{d}{}}{\mathrm{d}{\mathbf{{D}}}}\log\frac{\partial{\psi_{k}}}{\partial{% {\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \check{x}_{k}}}}(\bm{d}_{k}^{\text{T}}{\bm{Y}})}\right]}-{\mathbf{{D}}}^{-% \text{T}}.

That is, InfoMax ICA can be implemented as density estimation in a generative model with latent variables distributed independently and cumulatively according to 𝝍\bm{\psi} [5]; see Fig. 11.1.

a
b
Figure 11.1: “Infomax” independent-components analysis. (11.1a) The generative model. (11.1b) The discriminative model. Note that 𝐂=𝐃−1\mathbf{{C}}=\mathbf{{D}}^{-1}.

But we haven’t specified 𝝍\bm{\psi}! This omission may have seemed minor in the discriminative model—sigmoidal nonlinearities in neural networks are typically selected rather freely—but is striking in a generative model. And indeed, the choice matters. Suppose we had let the sigmoidal function be the CDF of a Gaussian. Then since we are modeling the observations as linear functions of the latent variables, 𝒀^=𝐂⁢𝑿^{\bm{\hat{Y}}}=\mathbf{{C}}{\bm{\hat{X}}}, their marginal distribution (Eq. 11.5) is clearly another mean-zero normal distribution, in particular 𝒩︀⁢(𝟎,𝐃T⁢𝐃−1)\mathcal{N}\mathopen{}\mathclose{{}\left(\bm{0},\>{\mathbf{{D}}^{\text{T}}% \mathbf{{D}}}^{-1}}\right)††margin: Exercise LABEL:ex: . Minimizing the loss in Eq. 11.4 then amounts merely to fitting the covariance of the observed data: 𝐃∗=𝔼⁢[𝒀⁢𝒀T]−1/2\mathbf{{D}}^{*}=\mathbb{E}{\mathopen{}\mathclose{{}\left[{\bm{Y}}{\bm{Y}}^{% \text{T}}}\right]}^{-1/2}. (This can also be shown by setting the gradient of Eq. 11.4, i.e. Eq. 7.29, to zero and solving for 𝐃\mathbf{{D}}.††margin: Exercise LABEL:ex: )

Figure 11.2: Sigmoidal functions. The logistic function (blue) and the Gaussian CDF (green).

If the observations are indeed normal, then whitening them in this way would indeed render them independent (since for jointly Gaussian random variables, uncorrelatedness implies independence)—but we do not need such an elaborate procedure to arrive at this conclusion! ICA is of interest precisely when the observations are not normal, in which case the optimal linear transformation cannot generally be stated a priori. Critically, squashing the data with the Gaussian CDF makes the outputs blind to the higher-order correlations, and is therefore not a suitable nonlinearity in cases of interest. In contrast, the (standard) logistic function is super-Gaussian (leptokurtotic), so InfoMax ICA with logistic outputs will generally do more than decorrelate its inputs. This may seem remarkable, given the visually minor discrepancy between the Gaussian CDF and the logistic function (Fig. 11.2; B.A. Olshausen, personal communication). Now we see the advantage of the generative perspective, from which this difference is more salient—and at long last, shed light on how to choose the feedforward nonlinearities, ψk\psi_{k}, in InfoMax ICA.

11.1.2 Nonlinear independent-component estimation

We now expand our view to nonlinear, albeit still invertible, transformations [11, 60, 12, 37]. In particular, consider a “generative function” 𝒄⁢(𝒙,𝜽)\bm{c}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}},\bm{\theta}) that consists of a series of invertible transformations. Once again, to emphasize that it is the inverse of a recognition or discriminative function, 𝒅⁢(𝒚,𝜽)\bm{d}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{y}},\bm{\theta}), we write 𝒄\bm{c} as 𝒅−1\bm{d}^{-1}:

equation (11.6) (11.6)
𝒚^=𝒄⁢(𝒙^,𝜽)=𝒅−1⁢(𝒙^,𝜽)=𝒅1−1∘𝒅2−1⁢⋯⁢𝒅L−1−1∘𝒅L−1⁢(𝒙^,𝜽).\bm{\hat{y}}=\bm{c}(\bm{\hat{x}},\bm{\theta})=\bm{d}^{-1}(\bm{\hat{x}},\bm{% \theta})=\bm{d}_{1}^{-1}\circ\bm{d}_{2}^{-1}\cdots\bm{d}_{{L}-1}^{-1}\circ\bm{% d}_{{L}}^{-1}(\bm{\hat{x}},\bm{\theta}).

This change of variables is called a flow††margin: flow [60]. Let us still assume a factorial prior, Eq. 11.2, and furthermore that it does not depend on any parameters. Since the transformations are (by assumption) invertible, the change-of-variables formula still applies. Therefore, Eq. 11.3 still holds, but the Jacobian determinant of composed functions becomes the product of the individual Jacobian determinants:

equation (11.7) (11.7)
p^⁢(𝒚;𝜽)=|∂𝝍∂𝒙^T⁢(𝒅⁢(𝒚,𝜽);𝜽)|⁢|∂𝒅∂𝒚T⁢(𝒚,𝜽)|=∏l=0L|𝐉l⁢(𝒚)|,{\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{};\bm{\theta}}\right)}=\mathopen% {}\mathclose{{}\left\lvert\frac{\partial{\bm{\psi}}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\hat{x}}}}^{% \text{T}}}\mathopen{}\mathclose{{}\left(\bm{d}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}},\bm{\theta});\bm{% \theta}}\right)}\right\rvert\mathopen{}\mathclose{{}\left\lvert\frac{\partial{% \bm{d}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{% rgb}{.75,0,.25}\bm{y}}}^{\text{T}}}({\color[rgb]{.75,0,.25}\definecolor[named]% {pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}},\bm{\theta})}\right\rvert\\ =\prod_{l=0}^{{L}}\mathopen{}\mathclose{{}\left\lvert\mathbf{J}_{l}({\color[% rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}})}% \right\rvert,

with Jacobians given by

𝐉0⁢(𝒚)=∂𝝍∂𝒙^T⁢(𝒅L∘⋯∘𝒅1⁢(𝒚,𝜽)),\displaystyle\mathbf{J}_{0}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}})=\frac{\partial{\bm{\psi}}}{\partial{{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \hat{x}}}}^{\text{T}}}\mathopen{}\mathclose{{}\left(\bm{d}_{{L}}\circ\cdots% \circ\bm{d}_{1}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb% }{.75,0,.25}\bm{y}},\bm{\theta})}\right),
𝐉l=∂𝒅l∂𝒙lT⁢(𝒅l−1∘⋯∘𝒅1⁢(𝒚,𝜽)).\displaystyle\mathbf{J}_{l}=\frac{\partial{\bm{d}_{l}}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}_{l}}}^{% \text{T}}}\mathopen{}\mathclose{{}\left(\bm{d}_{l-1}\circ\cdots\circ\bm{d}_{1}% ({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm% {y}},\bm{\theta})}\right).

(For the sake of writing derivatives, we have named the argument of the lthl^{\text{th}} function 𝒙l\bm{x}_{l}. This makes 𝒙1=𝒚\bm{x}_{1}=\bm{y}.) The functions induced by multiplying the initial distribution (the Jacobian determinant 𝐉0⁢(𝒚)\mathbf{J}_{0}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}% {.75,0,.25}\bm{y}}) at left) by, in turn, the determinants of each of the L{L} Jacobians 𝐉l\mathbf{J}_{l} at right, are “automatically” normalized and positive, and consequently valid probability distributions. This sequence is accordingly called a normalizing flow††margin: normalizing flow [60].

Since the generative function 𝒄⁢(𝒙,𝜽)=𝒅−1⁢(𝒙,𝜽)\bm{c}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}},\bm{\theta})=\bm{d}^{-1}({\color[rgb]{.75,0,.25}\definecolor% [named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}},\bm{\theta}) is invertible, we can certainly compute the arguments to the Jacobians. However, to keep the problem tractable, we also need to be able to compute the Jacobian determinants efficiently. Generically, this computation is cubic in the dimension of the data. This is intolerable, so we will generally limit the expressiveness of each 𝒄l\bm{c}_{l} to achieve something more practical.

Perhaps the most obvious limitation is to require that the transformations be “volume preserving”; that is, to require that Jacobian determinants are always unity [11] This can be achieved, for example, by splitting a data vector into two parts, and requiring (1) that the flow at a particular step ll of only one of the parts may depend on the other (this ensures that the Jacobian is upper triangular); and (2) that the flows of both parts depend on their previous values only through an identity transformation (this ensures that the two diagonal blocks of the Jacobian are identity matrices). In equations,

[𝒙l+1a𝒙l+1b]=[𝒙la𝒙lb+𝒎⁢(𝒙la,𝜽)]⟹∂𝒅l∂𝒙lT=[𝐈𝟎∂𝒎∂𝒙laT𝐈]⟹|∂𝒅l∂𝒙lT|=1.\begin{bmatrix}\bm{x}^{a}_{l+1}\\ \bm{x}^{b}_{l+1}\end{bmatrix}=\begin{bmatrix}\bm{x}^{a}_{l}\\ \bm{x}^{b}_{l}+\bm{m}(\bm{x}^{a}_{l},\bm{\theta})\end{bmatrix}\implies\frac{% \partial{\bm{d}_{l}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}_{l}}}^{\text{T}}}=\begin{bmatrix}\mathbf% {I}&\mathbf{0}\\ \frac{\partial{\bm{m}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}^{a}_{l}}}^{\text{T}}}&\mathbf{I}\end{% bmatrix}\implies\mathopen{}\mathclose{{}\left\lvert\frac{\partial{\bm{d}_{l}}}% {\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}_{l}}}^{\text{T}}}}\right\rvert=1.

… [[multiple layers of this]]

Now our loss is, as usual, the relative entropy. With the “recognition functions” 𝒅l\bm{d}_{l} of Eq. 11.6 and the corresponding model density of Eq. 11.7, the relative entropy becomes

equation (11.8) (11.8)
DKL⁡{p⁢(𝒀)∥∏l=0L|𝐉l⁢(𝒀)|}=𝔼𝒀⁢[log⁡p⁢(𝒀)−∑k=1Klog⁡(∂ψk∂x^k⁢(𝒀,𝜽))−∑l=1Llog⁡|∂𝒅l∂𝒙lT⁢(𝒀,𝜽)|].\operatorname*{\text{D}_{\text{KL}}}\mathopen{}\mathclose{{}\left\{{p\mathopen% {}\mathclose{{}\left({\bm{Y}}}\right)}\middle\|\prod_{l=0}^{{L}}\mathopen{}% \mathclose{{}\left\lvert\mathbf{J}_{l}({\bm{Y}})}\right\rvert}\right\}=\mathbb% {E}_{{\bm{Y}}{}}{\mathopen{}\mathclose{{}\left[\log{p\mathopen{}\mathclose{{}% \left({\bm{Y}}}\right)}-\sum_{k=1}^{{K}}\log\mathopen{}\mathclose{{}\left(% \frac{\partial{\psi_{k}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\hat{x}_{k}}}}({\bm{Y}},\bm{\theta})}\right)-% \sum_{l=1}^{{L}}\log\mathopen{}\mathclose{{}\left\lvert\frac{\partial{\bm{d}_{% l}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}_{l}}}^{\text{T}}}({\bm{Y}},\bm{\theta})}\right\rvert}\right]}.

(For concision, the Jacobians are written as a function directly of 𝒀{\bm{Y}}.)

The discriminative dual.

The model defined by Eq. 11.7, along with the loss in Eq. 11.8, has been called “nonlinear independent component analysis” (NICE) [11]. To see if the name is apposite, we employ our discriminative/generative duality, reinterpreting the minimization of the relative-entropy in Eq. 11.8 as a maximization of mutual information between the data, 𝒀{\bm{Y}}, and a random variable 𝒁^{\bm{\hat{Z}}} defined by the (reverse) flow in Eq. 11.6, 𝒅\bm{d}, followed by (elementwise) transformation by the CDF of the prior. Can this still be thought of as an “unmixing” operation, as in InfoMax ICA? The question is acute particularly in the case where the prior is chosen to be normal, since (as we have just seen) ICA reduces to whitening in such circumstances.

In this case, the generative marginal given by the normalizing flow, Eq. 11.7, becomes

p^⁢(𝒚;𝜽)=𝒩︀⁢(𝒅⁢(𝒚;𝜽);𝟎,𝐈)⁢|∂𝒅∂𝒚T⁢(𝒚,𝜽)|.{\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{};\bm{\theta}}\right)}=\mathcal{% N}\mathopen{}\mathclose{{}\left(\bm{d}({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}};\bm{\theta});\bm{0},\>\mathbf{I}% }\right)\mathopen{}\mathclose{{}\left\lvert\frac{\partial{\bm{d}}}{\partial{{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y% }}}^{\text{T}}}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb% }{.75,0,.25}\bm{y}},\bm{\theta})}\right\rvert.

Despite the appearance of a normal distribution in this expression, this marginal distribution is certainly not normal—even though the generative prior, p^𝑿^⁢(𝒙;𝜽){\hat{p}_{{\bm{\hat{X}}}}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{};\bm{\theta}}% \right)}, is. So fitting this p^⁢(𝒚;𝜽){\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{};\bm{\theta}}\right)} to the data will not in general merely fit their second-order statistics.

[[Connection to HMC]]