B.2 Probability and Statistics
Change of variables in probability densities
Distributions over linear combinations of random variables
-
•
Have:
-
•
Define:
-
•
Want:
Let us use the change-of-variables formula to eliminate in favor of . We will therefore need (1) to express as a function of the other variables,
and (2) the derivative of this function with respect to its dependence on the variable we are introducing, :
Then by the change-of-variables formula,
For the simple case of (“coordinate transformation”), this reduces to
Thus to get , one substitutes for in the original joint and integrates out .
The score function
The score is defined as the gradient of the log-likelihood (with respect to the parameters, ), . The mean of the score is zero:
The variance of the score is known as the Fisher information. Because its mean is zero, it is also the expected square of the score.
The Fisher information for exponential-family random variables
This turns out to take a simple form. For a (vector) random variable and “parameters” (that may themselves be random variables):
the Fisher information is:
where in the last line we have used the fact that the derivatives of the log-normalizer are the cumulants of the sufficient statistics () under the distribution. A perhaps more interesting equivalent can be derived by noting that:
Therefore, using the shorthand , we can write
Markov chains
Discrete random variables
[[[table]]]
Useful identities
Expectations of quadratic functions.
Consider a vector random variable with mean and covariance . Expectations are generally intractable for arbitrary functions of , but not low-order polynomials. In particular, expectations of first-order functions depend only on first-order expectations, i.e. the mean:
Similar, expectations of second-order functions depend on the first two moments. We can show this by exploiting the cyclic-permutation property of the matrix-trace operator, and the linearity of expectation and trace:
We can think of this as simply replacing each occurrence of the random variable in the original polynomial with its mean, plus a correction term to account for the covariance.
One very common application of this identity is to quadratic functions that take the form . This term can occur, for example, in the log probability of a Gaussian distribution about . Expanding this form to look like the polynomial in Eq. B.15, and then matching terms, we find
Simulating Poisson random variates with mean less than 1.
motivation…
Consider the graphical model shown below. We want to show that the marginal probabilitity of is distributed as a Poisson random variable with mean —as long as .
The derivation at right shows this marginalization. The third line follows because the probability of (the number of “successes”) is zero for any , since is the number of Bernoulli trials (it is impossible to have more successes than trials).