Skip to content.
An Introduction to ModernStatistical Learning
Preface
1
Introduction
I
Representation and Inference
II
Learning
III
Appendices
An Introduction to Modern
Statistical Learning
J.G. Makin
Contents
Preface
1
Introduction
1.1
Notation
1.1.1
Arguments and variables
1.1.2
Probabilistical functions and functionals
1.1.3
Derivatives
1.1.4
Other symbols
1.2
What Is a Generative Model?
1.2.1
Origin (1985)
1.2.2
Probabilistic graphical models (1990s)
1.2.3
Generative and recognition models (1990s)
1.2.4
Inference
1.2.5
The modern literature (2012 –)
I
Representation and Inference
2
Exponential Families
2.1
Maximum-Entropy Models
2.1.1
A second relationship between parameterizations
3
Directed Generative Models
3.1
Static models
3.1.1
Mixture models and the GMM
3.1.2
Jointly Gaussian models and factor analyzers
3.1.3
Sparse coding
3.2
Dynamical models
3.2.1
The hidden Markov Model and the
α
\alpha
-
γ
\gamma
algorithm
3.2.2
State-space models, the Kalman filter, and the RTS smoother
3.3
Exponential families
3.3.1
Posteriorization (Bayes inversion)
3.3.2
Marginalization
4
Undirected Generative Models
4.1
Undirected graphical models
4.1.1
Potential functions as unnormalized conditional distributions
4.2
The exponential-family harmonium
4.2.1
Derivation of the EFH
4.2.2
Generalized EFHs
4.2.3
Comparison with directed models
4.3
The Helmholtz machine
4.4
Recurrent EFHs
4.4.1
The recurrent temporal RBM
4.4.2
The recurrent EFH
5
General Algorithms for Exact Inference
5.1
Naïve inference
5.2
Variable Elimination (VE; the Elimination Algorithm)
5.3
Belief Propagation (the Sum-Product Algorithm)
5.3.1
Belief propagation on a clique tree.
5.3.2
The relationship between clique messages and sepatator potentials
5.4
The Junction-Tree Algorithm (JTA)
II
Learning
6
A Mathematical Framework for Learning
6.1
Learning as optimization
6.2
Information entropy
6.3
Fitting models to data
7
Learning Discriminative Models
7.1
Supervised learning
7.1.1
Linear regression
7.1.2
Generalized linear models
7.1.3
Artificial neural networks
7.2
Unsupervised learning
7.2.1
“InfoMax” in deterministic, invertible models
8
Learning Generative Models with Latent Variables
8.1
Introduction
8.2
Latent-variable density estimation
8.3
Expectation-Maximization
8.3.1
Derivation of EM
8.3.2
Information-theoretic perspectives on EM
9
Learning Invertible Generative Models
9.1
The Gaussian mixture model and
K
K
-means
9.1.1
K
K
-means
9.2
The hidden Markov model
9.3
Factor analysis and principal-components analysis
9.3.1
Principal-components analysis
9.4
Linear-Gaussian state-space models
10
Learning Non-Invertible Generative Models
10.1
Sparse coding with Gaussian recognition models
10.1.1
Sparse, independent priors and Gaussian emissions
10.1.2
Independent-components analysis
10.2
Variational Autoencoders
10.2.1
Monte Carlo gradient estimators
10.2.2
Examples
10.3
Diffusion Models
10.3.1
Gaussian diffusion models
10.4
Variational inference
11
Learning with Reparameterizations
11.1
A duality between generative and discriminative learning
11.1.1
InfoMax ICA, revisited
11.1.2
Nonlinear independent-component estimation
12
Learning Energy-Based Models
12.1
Direct minimization
12.1.1
Langevin dynamics
12.1.2
Contrastive divergence
12.2
The exponential-family harmonium
12.2.1
The EFH as a latent-variable model.
12.2.2
Learning with exponential-family harmoniums
12.2.3
Learning with deep belief networks
12.3
Contrastive losses
12.3.1
Noise-Contrastive Estimation
12.3.2
InfoNCE
12.3.3
“Local” NCE
III
Appendices
A
The Calculus of Variations and Its Applications
A.1
The Calculus of Variations
A.2
Applications of the Calculus of Variations
A.2.1
Lagrangian mechanics.
A.2.2
The method of adjoints
A.2.3
Pontryagin’s minimum principle
A.2.4
Neural ODEs
A.2.5
Hamiltonian Monte Carlo
A.2.6
Exponential families
B
Mathematical Appendix
B.1
Matrix Calculus
B.1.1
Derivatives with respect to vectors
B.1.2
Derivatives with respect to matrices
B.1.3
More useful identities
B.2
Probability and Statistics
B.3
Matrix Identities
C
Bonus Material
C.1
Leave-one-out cross validation for linear regression in one step
D
A Review of Probabilistic Graphical Models
Bibliography