Bibliography

  • [1] Azwar Abdulsalam and J.G. Makin. Revisiting Contrastive Divergence for Density Estimation and Sample Generation. Transaction on Machine Learning Research, October 2025.
  • [2] David H. Ackley, Geoffrey E. Hinton, and Terrence J. Sejnowski. A Learning Algorithm for Boltzmann Machines. Cognitive Science, 9(1):147–169, 1985.
  • [3] Anthony J. Bell and Terrence J. Sejnowski. An Information-Maximization Approach to Blind Separation and Blind Deconvolution. Neural Computation, 7(6):1129–1159, 1995.
  • [4] Christopher M Bishop. Pattern Recognition and Machine Learning. Springer, 2006.
  • [5] Jean-François Cardoso. Infomax and maximum likelihood for blind source separation. IEEE Signal Processing Letters, 4(4):112–114, 1997.
  • [6] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A Simple Framework for Contrastive Learning of Visual Representations. In ICML 2020, 2020.
  • [7] A. P. Dawid. Properties of Diagnostic Data Distributions. Biometrics, 32(3):647, 1976.
  • [8] Peter Dayan and L.F. Abbott. Theoretical Neuroscience. The MIT Press, 2005.
  • [9] Peter Dayan, Geoffrey E. Hinton, Radford M. Neal, and Richard S. Zemel. The Helmholtz Machine. Neural Computation, 7(5):889–904, 1995.
  • [10] A. P. Dempster, N. M. Laird, and D. B. Rubin. Maximum Likelihood from Incomplete Data Via the EM Algorithm. Journal of the Royal Statistical Society Series B: Statistical Methodology, 39(1):1–22, 1977.
  • [11] Laurent Dinh, David Krueger, and Yoshua Bengio. NICE: Non-linear Independent Components Estimation. In ICLR 2015 Workshop, 2015.
  • [12] Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using Real NVP. In International Conference on Learning Representations (ICLR) 2017, 2017.
  • [13] Yilun Du and Igor Mordatch. Implicit Generation and Generalization in Energy-Based Models. In Advances in Neural Information Processing Systems, 2019.
  • [14] M. Eden. Handwriting and pattern recognition. IEEE Transactions on Information Theory, 8(2):160–166, 1962.
  • [15] Bradley Efron. The Efficiency of Logistic Regression Compared to Normal Discriminant Analysis. Journal of the American Statistical Association, 70(352):892–898, 1975.
  • [16] R.A. Fisher. On the Mathematical Foundations of Theoretical Statistics. Philosophical Transactions of the Royal Society of London, Series A, 222:309–368, 1922.
  • [17] Stuart Geman and Donald Geman. Stochastic Relaxation, Gibbs Distributions, and the Bayesian Restoration of Images. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI-6(6):721–741, 1984.
  • [18] Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep Learning. MIT Press, 2016. http://www.deeplearningbook.org.
  • [19] Michael Gutmann and Aapo Hyvärinen. Noise-Contrastive Estimation of Unnormalized Statistical Models, with Applications to Natural Image Statistics. Journal of Machine Learning Research, 13:307–361, 2012.
  • [20] M. Halle and K. Stevens. Speech recognition: A model and a program for research. IEEE Transactions on Information Theory, 8(2):155–159, 1962.
  • [21] Per Christian Hansen, James G. Nagy, and Dianne P. O’Leary. Deblurring Images: Matrices, Spectra, and Filtering. SIAM, 2006.
  • [22] Michael Hartl. The Tau Manifesto. Accessed: 2022-05-09.
  • [23] John Hertz, Anders Krogh, Richard Palmer, and Roderick V. Jensen. Introduction to the Theory of Neural Computation. CRC Press, 2018.
  • [24] Geoffrey Hinton. Where Do Features Come From? Cognitive Science, 38(6):1078–1101, 2014.
  • [25] Geoffrey E. Hinton. Training Products of Experts by Minimizing Contrastive Divergence. Neural Computation, 14(8):1771–1800, 2002.
  • [26] Geoffrey E. Hinton. What kind of a graphical model is the brain. pages 1765–1775, 2005.
  • [27] Geoffrey E. Hinton and Andrew D. Brown. Spiking Boltzmann Machines. In Advances in Neural Information Processing Systems. MIT Press, 2000.
  • [28] Geoffrey E. Hinton and Zoubin Ghahramani. Generative models for discovering sparse distributed representations. Philosophical Transactions of the Royal Society B: Biological Sciences, 352(1358):1177–1190, 1997.
  • [29] Geoffrey E. Hinton and Richard S. Zemel. Autoencoders, Minimum Description Length and Helmholtz Free Energy. In Advances in Neural Information Processing Systems. Morgan Kaufmann, 1994.
  • [30] Matthew D. Hoffman and Matthew J. Johnson. ELBO Surgery: Yet Another Way to Carve Up the Variational Evidence Lower Bound. In Advances in Neural Information Processing Systems, 2016.
  • [31] Aapo Hyvärinen. Estimation of Non-Normalized Statistical Models by Score Matching. Journal of Machine Learning Research, 6:695–709, 2005.
  • [32] Eric Jang, Shixiang Gu, and Ben Poole. Categorical Reparameterization with Gumbel-Softmax. In ICLR 2017 Poster, 2017.
  • [33] E. T. Jaynes. Probability Theory. Cambridge University Press, 2003.
  • [34] Michael I. Jordan. Why the Logistic Function? A Tutorial Discussion on Probabilities and Neural Networks. Technical report, MIT Computational Cognitive Science Technical Report 9503, 1995.
  • [35] Michael I. Jordan. An Introduction to Probabilistic Graphical Models. Unpublished textbook, 2003.
  • [36] Diederik Kingma and Ruiqi Gao. Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation. In Advances in Neural Information Processing Systems. Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2023.
  • [37] Diederik P. Kingma and Prafulla Dhariwal. Glow: Generative Flow with Invertible 1x1 Convolutions. In Advances in Neural Information Processing Systems, 2018.
  • [38] Diederik P. Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. Variational Diffusion Models. In Advances in Neural Information Processing Systems, 2021.
  • [39] Diederik P Kingma and Max Welling. Auto-Encoding Variational Bayes. arXiv preprint, 2014.
  • [40] Friso H. Kingma, Pieter Abbeel, and Jonathan Ho. Bit-Swap: Recursive Bits-Back Coding for Lossless Compression with Hierarchical Latent Variables. In Proceedings of the 36th International Conference on Machine Learning (ICML 2019), 2019.
  • [41] S. L. Lauritzen and D. J. Spiegelhalter. Local Computations with Probabilities on Graphical Structures and Their Application to Expert Systems. Journal of the Royal Statistical Society Series B: Statistical Methodology, 50(2):157–194, 1988.
  • [42] Michael S. Lewicki and Terrence J. Sejnowski. Probabilistic Framework for the Adaptation and Comparison of Image Codes. Neural Computation, 11(7):1489–1517, 1999.
  • [43] Michael S. Lewicki and Terrence J. Sejnowski. Learning Overcomplete Representations. Neural Computation, 12(2):337–365, 2000.
  • [44] D. M. MACKAY. MINDLIKE BEHAVIOUR IN ARTEFACTS. The British Journal for the Philosophy of Science, 2(6):105–121, 1951.
  • [45] Chris J. Maddison, Andriy Mnih, and Yee Whye Teh. The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables. In International Conference on Learning Representations (ICLR 2017), 2017.
  • [46] McCulloch. UNMATCHED: A Logical Calculus of Ideas Immanent in Nervous Activity, 1943.
  • [47] Shakir Mohamed, Mihaela Rosca, Michael Figurnov, and Andriy Mnih. Monte Carlo Gradient Estimation in Machine Learning. Journal of Machine Learning Research, 21(132):1–62, 2020.
  • [48] Radford M. Neal and Geoffrey E. Hinton. A View of the Em Algorithm that Justifies Incremental, Sparse, and other Variants. In Learning in Graphical Models. Springer Netherlands, 1998.
  • [49] Ulric Neisser. Cognitive Psychology. Psychology Press, 2014.
  • [50] Erik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu, and Ying Nian Wu. On the Anatomy of MCMC-Based Maximum Likelihood Learning of Energy-Based Models. Proceedings of the AAAI Conference on Artificial Intelligence, 34(04):5272–5280, 2020.
  • [51] Erik Nijkamp, Mitch Hill, Song-Chun Zhu, and Ying Nian Wu. Learning Non-Convergent Non-Persistent Short-Run MCMC Toward Energy-Based Model. In Advances in Neural Information Processing Systems, 2019.
  • [52] Bruno A. Olshausen and David J. Field. Sparse coding with an overcomplete basis set: A strategy employed by V1? Vision Research, 37(23):3311–3325, 1997.
  • [53] Manfred Opper and Cédric Archambeau. The Variational Gaussian Approximation Revisited. Neural Computation, 21(3):786–792, 2009.
  • [54] John Paisley, David Blei, and Michael Jordan. Variational Bayesian Inference with Stochastic Search. In Proceedings of the 29th International Conference on Machine Learning (ICML 2012), 2012.
  • [55] Judea Pearl. Reverend Bayes on Inference Engines: A Distributed Hierarchical Approach. In AAAI 1982, 1982.
  • [56] Arthur E.C. Pece. The Problem of Sparse Image Coding. Journal of Mathematical Imaging and Vision, 17(2):89–108, 2002.
  • [57] L. R. Pericchi and A. F. M. Smith. Exact and Approximate Posterior Moments for a Normal Location Parameter. Journal of the Royal Statistical Society Series B: Statistical Methodology, 54(3):793–804, 1992.
  • [58] W.V.O. Quine. Two Dogmas of Empiricism. The Philosophical Review, 60:20–43, 1951.
  • [59] Michael Revow, Christopher K. I. Williams, and Geoffrey E. Hinton. Using Generative Models for Handwritten Digit Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 18(6):592–606, 1996.
  • [60] Danilo Jimenez Rezende and Shakir Mohamed. Variational Inference with Normalizing Flows. In ICML 2015, 2015.
  • [61] Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. Stochastic Backpropagation and Approximate Inference in Deep Generative Models. In ICML 2014, 2014.
  • [62] Herbert Robbins. AN EMPIRICAL BAYES APPROACH TO STATISTICS. In Contribution to the Theory of Statistics. University of California Press, 1956.
  • [63] Sam Roweis and Zoubin Ghahramani. A Unifying Review of Linear Gaussian Models. Neural Computation, 11(2):305–345, 1999.
  • [64] Y.D. Rubinstein and Trevor Hastie. Discriminative vs Informative Learning. KDD-97 Proceedings, pages 49–53, 1997.
  • [65] Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. wav2vec: Unsupervised Pre-training for Speech Recognition. In INTERSPEECH 2019, 2019.
  • [66] C. E. Shannon. A Mathematical Theory of Communication. Bell System Technical Journal, 27(4):623–656, 1948.
  • [67] Eero P Simoncelli and Bruno A Olshausen. Natural Image Statistics and Neural Representation. Annual Review of Neuroscience, 24(1):1193–1216, 2001.
  • [68] Paul Smolensky. Information Processing in Dynamical Systems: Foundations of Harmony Theory. In Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Vol. 1. MIT Press, 1986.
  • [69] Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep Unsupervised Learning Using Nonequilibrium Thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning (ICML 2015), 2015.
  • [70] Richard M. Soland. Bayesian Analysis of the Weibull Process With Unknown Scale and Shape Parameters. IEEE Transactions on Reliability, R-18(4):181–184, 1969.
  • [71] Yang Song and Stefano Ermon. Generative Modeling by Estimating Gradients of the Data Distribution. In Advances in Neural Information Processing Systems, 2019.
  • [72] Michael E. Tipping and Christopher M. Bishop. Probabilistic Principal Component Analysis. Journal of the Royal Statistical Society Series B: Statistical Methodology, 61(3):611–622, 1999.
  • [73] D.M. Titterington, A.F.M. Smith, and U.E. Makov. Statistical Analysis of Finite Mixture Distributions. Wiley, 1985.
  • [74] James Townsend, Thomas Bird, and David Barber. Practical Lossless Compression with Latent Variables Using Bits Back Coding. In International Conference on Learning Representations (ICLR 2019), 2019.
  • [75] Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation Learning with Contrastive Predictive Coding, 2018.
  • [76] Pascal Vincent. A Connection Between Score Matching and Denoising Autoencoders. Neural Computation, 23(7):1661–1674, 2011.
  • [77] Hermann von Helmholtz. Helmholtz’s Treatise on Physiological Optics, volume 1–3. Optical Society of America, Rochester, NY, 3rd german edition, 1924–1925. Translated from the 3rd German ed. (1909–1911).
  • [78] Max Welling, Michal Rosen-Zvi, and Geoffrey E. Hinton. Exponential Family Harmoniums with an Application to Information Retrieval. In Advances in Neural Information Processing Systems. MIT Press, 2004.