Skip to content.
Bibliography
-
[1]
Azwar Abdulsalam and J.G. Makin.
Revisiting Contrastive Divergence for Density Estimation and Sample
Generation.
Transaction on Machine Learning Research, October 2025.
-
[2]
David H. Ackley, Geoffrey E. Hinton, and Terrence J. Sejnowski.
A Learning Algorithm for Boltzmann Machines.
Cognitive Science, 9(1):147–169, 1985.
-
[3]
Anthony J. Bell and Terrence J. Sejnowski.
An Information-Maximization Approach to Blind Separation and Blind
Deconvolution.
Neural Computation, 7(6):1129–1159, 1995.
-
[4]
Christopher M Bishop.
Pattern Recognition and Machine Learning.
Springer, 2006.
-
[5]
Jean-François Cardoso.
Infomax and maximum likelihood for blind source separation.
IEEE Signal Processing Letters, 4(4):112–114, 1997.
-
[6]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton.
A Simple Framework for Contrastive Learning of Visual
Representations.
In ICML 2020, 2020.
-
[7]
A. P. Dawid.
Properties of Diagnostic Data Distributions.
Biometrics, 32(3):647, 1976.
-
[8]
Peter Dayan and L.F. Abbott.
Theoretical Neuroscience.
The MIT Press, 2005.
-
[9]
Peter Dayan, Geoffrey E. Hinton, Radford M. Neal, and Richard S. Zemel.
The Helmholtz Machine.
Neural Computation, 7(5):889–904, 1995.
-
[10]
A. P. Dempster, N. M. Laird, and D. B. Rubin.
Maximum Likelihood from Incomplete Data Via the EM Algorithm.
Journal of the Royal Statistical Society Series B: Statistical
Methodology, 39(1):1–22, 1977.
-
[11]
Laurent Dinh, David Krueger, and Yoshua Bengio.
NICE: Non-linear Independent Components Estimation.
In ICLR 2015 Workshop, 2015.
-
[12]
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio.
Density estimation using Real NVP.
In International Conference on Learning Representations (ICLR)
2017, 2017.
-
[13]
Yilun Du and Igor Mordatch.
Implicit Generation and Generalization in Energy-Based Models.
In Advances in Neural Information Processing Systems, 2019.
-
[14]
M. Eden.
Handwriting and pattern recognition.
IEEE Transactions on Information Theory, 8(2):160–166, 1962.
-
[15]
Bradley Efron.
The Efficiency of Logistic Regression Compared to Normal
Discriminant Analysis.
Journal of the American Statistical Association,
70(352):892–898, 1975.
-
[16]
R.A. Fisher.
On the Mathematical Foundations of Theoretical Statistics.
Philosophical Transactions of the Royal Society of London,
Series A, 222:309–368, 1922.
-
[17]
Stuart Geman and Donald Geman.
Stochastic Relaxation, Gibbs Distributions, and the Bayesian
Restoration of Images.
IEEE Transactions on Pattern Analysis and Machine Intelligence,
PAMI-6(6):721–741, 1984.
-
[18]
Ian Goodfellow, Yoshua Bengio, and Aaron Courville.
Deep Learning.
MIT Press, 2016.
http://www.deeplearningbook.org.
-
[19]
Michael Gutmann and Aapo Hyvärinen.
Noise-Contrastive Estimation of Unnormalized Statistical Models,
with Applications to Natural Image Statistics.
Journal of Machine Learning Research, 13:307–361, 2012.
-
[20]
M. Halle and K. Stevens.
Speech recognition: A model and a program for research.
IEEE Transactions on Information Theory, 8(2):155–159, 1962.
-
[21]
Per Christian Hansen, James G. Nagy, and Dianne P. O’Leary.
Deblurring Images: Matrices, Spectra, and Filtering.
SIAM, 2006.
-
[22]
Michael Hartl.
The Tau Manifesto.
Accessed: 2022-05-09.
-
[23]
John Hertz, Anders Krogh, Richard Palmer, and Roderick V. Jensen.
Introduction to the Theory of Neural Computation.
CRC Press, 2018.
-
[24]
Geoffrey Hinton.
Where Do Features Come From?
Cognitive Science, 38(6):1078–1101, 2014.
-
[25]
Geoffrey E. Hinton.
Training Products of Experts by Minimizing Contrastive Divergence.
Neural Computation, 14(8):1771–1800, 2002.
-
[26]
Geoffrey E. Hinton.
What kind of a graphical model is the brain.
pages 1765–1775, 2005.
-
[27]
Geoffrey E. Hinton and Andrew D. Brown.
Spiking Boltzmann Machines.
In Advances in Neural Information Processing Systems. MIT
Press, 2000.
-
[28]
Geoffrey E. Hinton and Zoubin Ghahramani.
Generative models for discovering sparse distributed
representations.
Philosophical Transactions of the Royal Society B: Biological
Sciences, 352(1358):1177–1190, 1997.
-
[29]
Geoffrey E. Hinton and Richard S. Zemel.
Autoencoders, Minimum Description Length and Helmholtz Free Energy.
In Advances in Neural Information Processing Systems. Morgan
Kaufmann, 1994.
-
[30]
Matthew D. Hoffman and Matthew J. Johnson.
ELBO Surgery: Yet Another Way to Carve Up the Variational Evidence
Lower Bound.
In Advances in Neural Information Processing Systems, 2016.
-
[31]
Aapo Hyvärinen.
Estimation of Non-Normalized Statistical Models by Score Matching.
Journal of Machine Learning Research, 6:695–709, 2005.
-
[32]
Eric Jang, Shixiang Gu, and Ben Poole.
Categorical Reparameterization with Gumbel-Softmax.
In ICLR 2017 Poster, 2017.
-
[33]
E. T. Jaynes.
Probability Theory.
Cambridge University Press, 2003.
-
[34]
Michael I. Jordan.
Why the Logistic Function? A Tutorial Discussion on Probabilities
and Neural Networks.
Technical report, MIT Computational Cognitive Science Technical
Report 9503, 1995.
-
[35]
Michael I. Jordan.
An Introduction to Probabilistic Graphical Models.
Unpublished textbook, 2003.
-
[36]
Diederik Kingma and Ruiqi Gao.
Understanding Diffusion Objectives as the ELBO with Simple Data
Augmentation.
In Advances in Neural Information Processing Systems. Neural
Information Processing Systems Foundation, Inc. (NeurIPS), 2023.
-
[37]
Diederik P. Kingma and Prafulla Dhariwal.
Glow: Generative Flow with Invertible 1x1 Convolutions.
In Advances in Neural Information Processing Systems, 2018.
-
[38]
Diederik P. Kingma, Tim Salimans, Ben Poole, and Jonathan Ho.
Variational Diffusion Models.
In Advances in Neural Information Processing Systems, 2021.
-
[39]
Diederik P Kingma and Max Welling.
Auto-Encoding Variational Bayes.
arXiv preprint, 2014.
-
[40]
Friso H. Kingma, Pieter Abbeel, and Jonathan Ho.
Bit-Swap: Recursive Bits-Back Coding for Lossless Compression with
Hierarchical Latent Variables.
In Proceedings of the 36th International Conference on Machine
Learning (ICML 2019), 2019.
-
[41]
S. L. Lauritzen and D. J. Spiegelhalter.
Local Computations with Probabilities on Graphical Structures and
Their Application to Expert Systems.
Journal of the Royal Statistical Society Series B: Statistical
Methodology, 50(2):157–194, 1988.
-
[42]
Michael S. Lewicki and Terrence J. Sejnowski.
Probabilistic Framework for the Adaptation and Comparison of Image
Codes.
Neural Computation, 11(7):1489–1517, 1999.
-
[43]
Michael S. Lewicki and Terrence J. Sejnowski.
Learning Overcomplete Representations.
Neural Computation, 12(2):337–365, 2000.
-
[44]
D. M. MACKAY.
MINDLIKE BEHAVIOUR IN ARTEFACTS.
The British Journal for the Philosophy of Science,
2(6):105–121, 1951.
-
[45]
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh.
The Concrete Distribution: A Continuous Relaxation of Discrete
Random Variables.
In International Conference on Learning Representations (ICLR
2017), 2017.
-
[46]
McCulloch.
UNMATCHED: A Logical Calculus of Ideas Immanent in Nervous
Activity, 1943.
-
[47]
Shakir Mohamed, Mihaela Rosca, Michael Figurnov, and Andriy Mnih.
Monte Carlo Gradient Estimation in Machine Learning.
Journal of Machine Learning Research, 21(132):1–62, 2020.
-
[48]
Radford M. Neal and Geoffrey E. Hinton.
A View of the Em Algorithm that Justifies Incremental, Sparse, and
other Variants.
In Learning in Graphical Models. Springer Netherlands, 1998.
-
[49]
Ulric Neisser.
Cognitive Psychology.
Psychology Press, 2014.
-
[50]
Erik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu, and Ying Nian Wu.
On the Anatomy of MCMC-Based Maximum Likelihood Learning of
Energy-Based Models.
Proceedings of the AAAI Conference on Artificial Intelligence,
34(04):5272–5280, 2020.
-
[51]
Erik Nijkamp, Mitch Hill, Song-Chun Zhu, and Ying Nian Wu.
Learning Non-Convergent Non-Persistent Short-Run MCMC Toward
Energy-Based Model.
In Advances in Neural Information Processing Systems, 2019.
-
[52]
Bruno A. Olshausen and David J. Field.
Sparse coding with an overcomplete basis set: A strategy employed by
V1?
Vision Research, 37(23):3311–3325, 1997.
-
[53]
Manfred Opper and Cédric Archambeau.
The Variational Gaussian Approximation Revisited.
Neural Computation, 21(3):786–792, 2009.
-
[54]
John Paisley, David Blei, and Michael Jordan.
Variational Bayesian Inference with Stochastic Search.
In Proceedings of the 29th International Conference on Machine
Learning (ICML 2012), 2012.
-
[55]
Judea Pearl.
Reverend Bayes on Inference Engines: A Distributed Hierarchical
Approach.
In AAAI 1982, 1982.
-
[56]
Arthur E.C. Pece.
The Problem of Sparse Image Coding.
Journal of Mathematical Imaging and Vision, 17(2):89–108,
2002.
-
[57]
L. R. Pericchi and A. F. M. Smith.
Exact and Approximate Posterior Moments for a Normal Location
Parameter.
Journal of the Royal Statistical Society Series B: Statistical
Methodology, 54(3):793–804, 1992.
-
[58]
W.V.O. Quine.
Two Dogmas of Empiricism.
The Philosophical Review, 60:20–43, 1951.
-
[59]
Michael Revow, Christopher K. I. Williams, and Geoffrey E. Hinton.
Using Generative Models for Handwritten Digit Recognition.
IEEE Transactions on Pattern Analysis and Machine Intelligence,
18(6):592–606, 1996.
-
[60]
Danilo Jimenez Rezende and Shakir Mohamed.
Variational Inference with Normalizing Flows.
In ICML 2015, 2015.
-
[61]
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra.
Stochastic Backpropagation and Approximate Inference in Deep
Generative Models.
In ICML 2014, 2014.
-
[62]
Herbert Robbins.
AN EMPIRICAL BAYES APPROACH TO STATISTICS.
In Contribution to the Theory of Statistics. University of
California Press, 1956.
-
[63]
Sam Roweis and Zoubin Ghahramani.
A Unifying Review of Linear Gaussian Models.
Neural Computation, 11(2):305–345, 1999.
-
[64]
Y.D. Rubinstein and Trevor Hastie.
Discriminative vs Informative Learning.
KDD-97 Proceedings, pages 49–53, 1997.
-
[65]
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli.
wav2vec: Unsupervised Pre-training for Speech Recognition.
In INTERSPEECH 2019, 2019.
-
[66]
C. E. Shannon.
A Mathematical Theory of Communication.
Bell System Technical Journal, 27(4):623–656, 1948.
-
[67]
Eero P Simoncelli and Bruno A Olshausen.
Natural Image Statistics and Neural Representation.
Annual Review of Neuroscience, 24(1):1193–1216, 2001.
-
[68]
Paul Smolensky.
Information Processing in Dynamical Systems: Foundations of Harmony
Theory.
In Parallel Distributed Processing: Explorations in the
Microstructure of Cognition, Vol. 1. MIT Press, 1986.
-
[69]
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli.
Deep Unsupervised Learning Using Nonequilibrium Thermodynamics.
In Proceedings of the 32nd International Conference on Machine
Learning (ICML 2015), 2015.
-
[70]
Richard M. Soland.
Bayesian Analysis of the Weibull Process With Unknown Scale and
Shape Parameters.
IEEE Transactions on Reliability, R-18(4):181–184, 1969.
-
[71]
Yang Song and Stefano Ermon.
Generative Modeling by Estimating Gradients of the Data
Distribution.
In Advances in Neural Information Processing Systems, 2019.
-
[72]
Michael E. Tipping and Christopher M. Bishop.
Probabilistic Principal Component Analysis.
Journal of the Royal Statistical Society Series B: Statistical
Methodology, 61(3):611–622, 1999.
-
[73]
D.M. Titterington, A.F.M. Smith, and U.E. Makov.
Statistical Analysis of Finite Mixture Distributions.
Wiley, 1985.
-
[74]
James Townsend, Thomas Bird, and David Barber.
Practical Lossless Compression with Latent Variables Using Bits Back
Coding.
In International Conference on Learning Representations (ICLR
2019), 2019.
-
[75]
Aaron van den Oord, Yazhe Li, and Oriol Vinyals.
Representation Learning with Contrastive Predictive Coding, 2018.
-
[76]
Pascal Vincent.
A Connection Between Score Matching and Denoising Autoencoders.
Neural Computation, 23(7):1661–1674, 2011.
-
[77]
Hermann von Helmholtz.
Helmholtz’s Treatise on Physiological Optics, volume 1–3.
Optical Society of America, Rochester, NY, 3rd german edition,
1924–1925.
Translated from the 3rd German ed. (1909–1911).
-
[78]
Max Welling, Michal Rosen-Zvi, and Geoffrey E. Hinton.
Exponential Family Harmoniums with an Application to Information
Retrieval.
In Advances in Neural Information Processing Systems. MIT
Press, 2004.