6.3 Fitting models to data
The reader has probably anticipated our approach to fitting models to data: in broad terms, we will attempt to minimize the relative entropy of the data distribution and a model distribution. Most (but not all) of the losses in this book, then, whether for discriminative or generative models, supervised or unsupervised learning, will be written as relative entropies, with the model distribution in place of . This measures the number of extra bits that we would have to use to encode the data if we were to use the model rather than the (unattainable) data distribution to construct the encoding scheme.
There are, however, some subtleties, chiefly arising from the fact that we have access only to samples from the data distribution. Several approaches are possible here and appear in the literature. Perhaps the most common is to assume that the data are independent and identically distributed (i.i.d.). We make the same assumption for the model. Then we let the optimal parameters be those that minimize the relative entropy of the marginal distributions,
Thus we see that the same parameters minimize cross and relative entropy, and that they are the maximum-likelihood estimate††margin: maximum-likelihood estimate .
However, this approach is not wholly satisfactory, since it holds only for i.i.d. data. The i.i.d. assumption is not per se unsatisfactory. The problem is that we would prefer to embed all of our assumptions into the model; or, more felicitously, to make only modeling decisions, not assumptions. This allows us to separate cleanly the aspects of the problem over which we do and do not have control…..
Maximum-likelihood estimates (MLEs) have long enjoyed widespread use in statistics for their asymptotic properties: Suppose the data were “actually” generated from a parameterized distribution within the family of model distributions, . Then as the sample size approaches infinity, the MLE converges (in probability) to the true underlying parameter (“consistency”) and achieves the mininum mean squared error among all consistent estimators (“efficiency”). The parameter estimates of models fit by minimizing relative entropy inherit these properties.
[[Forward and reverse KL]]
[[continuous RVs]]
[[well known losses like squared error and the “binary cross entropy”]]