Chapter 12 Learning Energy-Based Models

One of the basic problems we have been grappling with in fitting generative models to data is how to make the model sufficiently expressive. For example, some of the complexity or “lumpiness” of the data distribution can be explained as the effect of marginalizing out some latent variables—as in a mixture of Gaussians. As we have seen, GMMs are not sufficient to model (e.g.) natural images, so we need to introduce more complexity. VAEs can be thought of as generalizations of GMMs11 1 Or, better, factor analyzers, in which the latent variable provides the input to a continuous, highly nonlinear, mean function, rather than merely the index into a finite set of means. Normalizing flows likewise use complex nonlinearities (neural networks) to map simply-distributed latent variables into variables with more complicated distributions, although they treat the output of the neural network itself as the random variable of interest (they don’t bother to add a little Gaussian noise)—but at the price that the network must be invertible.

An alternative to these is to model the unnormalized distribution of observed variables, or equivalently, the energy, which defines the probability by way of a Boltzmann††margin: Boltzmann distribution or Gibbs distribution††margin: Gibbs distribution 22 2 These terms have more, and subtly different, content in statistical physics, but in machine learning they are essentially synonyms for Eq. 12.1:

equation (12.1) (12.1)
U^(𝒚,𝜽) . . =−logp^(𝒚;𝜽)−logZ(𝜽)⟹p^(𝒚;𝜽)=1Z⁢(𝜽)exp{−U^(𝒚,𝜽)}.\hat{U}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[named% ]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{},\bm{\theta}}\right)\mathrel{\vbox{% \hbox{.}\hbox{.} }}=-\log{\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{};\bm{\theta}}% \right)}-\log Z(\bm{\theta})\implies{\hat{p}\mathopen{}\mathclose{{}\left({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y% }}{};\bm{\theta}}\right)}=\frac{1}{Z(\bm{\theta})}\exp\mathopen{}\mathclose{{}% \left\{-\hat{U}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{},\bm{\theta}}% \right)}\right\}.

The advantage of such an energy-based model (EBM)††margin: energy-based model (EBM) is that U^⁢(𝒚,𝜽)\hat{U}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[named% ]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{},\bm{\theta}}\right) can be an arbitrarily complex function mapping to the real line and, consequently, we are not limited to distributions with a known parametric form (like Gaussian or Poisson or etc.), or that can be constructed out of invertible transformations of noise. The seemingly fatal disadvantages are the following:

  1. 1.
    ​

    Without a parametric model, it is not immediately obvious how to generate samples.

  2. 2.
    ​

    Latent variables are appealing not merely for the complexity they can add to the marginal distribution, but for providing more useful representations of the data. So we have lost something in doing without latent variables.

  3. 3.
    ​

    Computing the normalizer will be intractable if we let U^⁢(𝒚,𝜽)\hat{U}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[named% ]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{y}}{},\bm{\theta}}\right) be particularly complex—which was the whole point of taking this approach! Without the normalizer, we cannot assign a probability to any datum (although we can assign relative probabilities to any pair of data); nor is it obvious how to fit such a model to data.