A.2 Applications of the Calculus of Variations

A.2.1 Lagrangian mechanics.

The Lagrangian for mechanical systems.

The motivation for Lagrangian mechanics is constrained systems. We could solve each constraint for a variable and then eliminate that variable from the equations of motion. But this is often onerous, so instead we use Lagrange multipliers. What is the loss function to which we add these Lagrange-multiplier terms? For typical mechanical systems, it is the difference of kinetic and potential energies:

ℓ(q,q˙,t) . . =K(q˙)−U(q)=12mq˙2−U(q);\ell({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25% }q},{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \dot{q}},t)\mathrel{\vbox{\hbox{.}\hbox{.} }}=K({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25% }\dot{q}})-U({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}q})=\frac{1}{2}m{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\dot{q}}^{2}-U({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}q});

or, for vector-valued coordinates,11 1 We use 𝒒\bm{q} for the coordinates to remind ourselves that they are generalized (i.e., not necessarily Cartesian).

equation (A.5) (A.5)
ℓ(𝒒,𝒒˙,t) . . =K(𝒒˙)−U(𝒒)=12𝒒˙(t)T𝐌𝒒˙(t)−U(𝒒(t)).\ell({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25% }\bm{q}},{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\dot{q}}},t)\mathrel{\vbox{\hbox{.}\hbox{.} }}=K({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25% }\bm{\dot{q}}})-U({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{% rgb}{.75,0,.25}\bm{q}})=\frac{1}{2}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}}(t)^{\text{T}}\mathbf{M}{\color[% rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}% }(t)-U({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{q}}(t)).

This cost accumulates over time. Ignoring for simplicity any constraints, i.e. letting ℒ︀=ℓ\mathcal{L}=\ell, the total cost is therefore

J⁢[𝒒]=∫0Tℓ⁢(𝒒⁢(t),𝒒˙⁢(t),t)⁢dt.J[{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{q}}]=\int_{0}^{T}\ell({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}}(t),{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}}(t),t)\mathop{}\!\mathrm{d}% {t{}}.

This is a functional, so to find the minimizing function of time, 𝒒∗⁢(t)\bm{q}^{*}(t), we solve the Euler-Lagrange equations, Eq. A.1. This yields

𝐌⁢𝒒¨∗=−∂U∂𝒒⁢(𝒒∗).\mathbf{M}\bm{\ddot{q}}^{*}=-\frac{\partial{U}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}}}}}(\bm{q}% ^{*}).

So e.g. if the potential energy is due to gravity (−m⁢g⁢q-mg{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}q}) and a spring force (k⁢q2/2k{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}q}^% {2}/2), the force on the RHS is

k⁢q−m⁢g,k{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}q}-mg,

so that we have

m⁢q¨∗+k⁢q∗=m⁢g,m\ddot{q}^{*}+kq^{*}=mg,

the usual equations of motion from Newtons’s laws.

Note, however, that we have derived a conservative system, in which energy is traded between potential and kinetics, but is never destroyed—there is no b⁢q˙b{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \dot{q}} term. We return to this point below.

The Hamiltonian.

Notice that the mechanical Lagrangian above is convex in the velocity, q˙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\dot% {q}}. This will be the case for most physical systems of interest. Consequently, ∂ℒ︀/∂𝒒˙\partial{\mathcal{L}}/\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}}}} will be an invertible function of 𝒒˙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \dot{q}}} for all values of 𝒒{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% q}} and tt,

equation (A.6) (A.6)
𝐠(𝒒˙) . . =∂ℒ︀∂𝒒˙(𝒒,𝒒˙,t)= . . 𝒑(t).\mathbf{g}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\dot{q}}})\mathrel{\vbox{\hbox{.}\hbox{.} }}=\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}}}}}(\bm{q},{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}},t)=% \mathrel{\vbox{\hbox{.}\hbox{.} }}{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{p}}(t).

The second equality follows from our definition of the conjugate variables, Eq. A.3. This suggests a certain change of variables, namely, substituting 𝐠−1⁢(𝒑){\mathbf{g}}^{-1}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{% rgb}{.75,0,.25}\bm{p}}) for all occurrences of 𝒒˙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \dot{q}}} in ℒ︀\mathcal{L}.

However, a more profound transform is in fact available.22 2 Profound, but infrequently useful: the Lagrangian formulation of physical problems is almost always easier to solve. But the Hamiltonian formulation is often more elegant. If we consider the Lagrangian itself as a function of the velocities, we note that it can equally well be represented by the ordered input/output pairs (𝒒˙,ℒ︀)(\bm{\dot{q}},\mathcal{L}) as by the pairs of slopes and intercepts of the tangent lines. The slopes are precisely the conjugate variables (Eq. A.6), 𝒑\bm{p}, and for the (negative) intercepts we introduce the symbol ℋ︀\mathcal{H}. So we can represent the Lagrangian with ordered pairs (𝒑,ℋ︀)(\bm{p},\mathcal{H}). Moving between these representations is known as a Legendre transformation. From simple manipulation of the equation of a tangent line††margin: Also include the more general definition in terms of infimum , we can get ℋ︀\mathcal{H} in terms of everything else:

equation (A.7) (A.7)
ℒ︀⁢(𝒒,𝒒˙,t)\displaystyle\mathcal{L}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}},{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}},t)
=∂ℒ︀∂𝒒˙T⁢𝒒˙−ℋ︀⁢(𝒒,𝐠⁢(𝒒˙),t)\displaystyle{}=\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}}}}^{\text{T}}}% {\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \dot{q}}}-\mathcal{H}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}},\mathbf{g}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}}),t)
y=m⁢x+b\displaystyle{}y=mx+b
⟹ℋ︀⁢(𝒒,𝒑,t)\displaystyle\implies\mathcal{H}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}},{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}},t)
=𝒑T⁢𝐠−1⁢(𝒑)−ℒ︀⁢(𝒒,𝐠−1⁢(𝒑),t)\displaystyle{}={\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb% }{.75,0,.25}\bm{p}}^{\text{T}}{\mathbf{g}}^{-1}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}})-\mathcal{L}({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{q% }},{\mathbf{g}}^{-1}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor% }{rgb}{.75,0,.25}\bm{p}}),t)
Eqs. A.3 and A.6.

For mechanical systems in Cartesian coordinates, the (𝒑,ℋ︀)(\bm{p},\mathcal{H}) representation has an intuitive interpretation. Applying the definition of the conjugate variables (Eq. A.6) to the mechanical Lagrangian in Eq. A.5, we obtain 𝒑=𝐌⁢𝒒˙{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% p}}=\mathbf{M}{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\dot{q}}}—the momenta of the system. The Hamiltonian, for its part, turns out to be the total energy of the system, the sum of the kinetic and potential energies:

ℋ︀⁢(𝒒,𝐠⁢(𝒒˙),t)=𝒒˙T⁢𝐌⁢𝒒˙−(K⁢(𝒒˙)−U⁢(𝒒))=2⁢K⁢(𝒒˙)−(K⁢(𝒒˙)−U⁢(𝒒))=K⁢(𝒒˙)+U⁢(𝒒)⟹ℋ︀⁢(𝒒,𝒑,t)=K⁢(𝐌−1⁢𝒑)+U⁢(𝒒)=12⁢𝒑⁢(t)T⁢𝐌−1⁢𝒑⁢(t)+U⁢(𝒒⁢(t)).\begin{split}\mathcal{H}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}},\mathbf{g}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}}),t)&{}={% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \dot{q}}}^{\text{T}}\mathbf{M}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}}-\mathopen{}\mathclose{{}\left(K({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \dot{q}}})-U({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{q}})}\right)\\ &{}=2K({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\dot{q}}})-\mathopen{}\mathclose{{}\left(K({\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}})-U(% {\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% q}})}\right)\\ &{}=K({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\dot{q}}})+U({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}})\\ \implies\mathcal{H}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}% {rgb}{.75,0,.25}\bm{q}},{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}},t)&{}=K({\mathbf{M}}^{-1}{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}})+U({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{q% }})\\ &{}=\frac{1}{2}{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}% {.75,0,.25}\bm{p}}(t)^{\text{T}}{\mathbf{M}}^{-1}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}(t)+U({\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}}(t)).\end{split}

The Euler-Lagrange equations in Hamiltonian mechanics.

We saw above the implications of the Euler-Lagrange equations for a mechanical Lagrangian—to wit, they give rise to Newton’s laws. Here we investigate their implications for the Hamiltonian.

∂ℋ︀∂𝒒T⁢(𝒒∗,𝒑∗,t)\displaystyle\frac{\partial{\mathcal{H}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}}}}^{\text{T}}}(\bm{q% }^{*},\bm{p}^{*},t)
=dd⁢𝒒∗T⁢(𝒑∗T⁢𝐠−1⁢(𝒑∗)−ℒ︀⁢(𝒒∗,𝐠−1⁢(𝒑∗),t))\displaystyle{}=\frac{\mathrm{d}{}}{\mathrm{d}{\bm{q}^{*}}^{\text{T}}}% \mathopen{}\mathclose{{}\left({\bm{p}^{*}}^{\text{T}}{\mathbf{g}}^{-1}(\bm{p}^% {*})-\mathcal{L}(\bm{q}^{*},{\mathbf{g}}^{-1}(\bm{p}^{*}),t)}\right)
Eq. A.7
=−𝒑˙∗T\displaystyle{}=-{\dot{\bm{p}}^{*}{}}^{\text{T}}
Eq. A.3
∂ℋ︀∂𝒑T⁢(𝒒,𝒑,t)\displaystyle\frac{\partial{\mathcal{H}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}}}^{\text{T}}}({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{q% }},{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{p}},t)
=dd⁢𝒑T⁢(𝒑T⁢𝐠−1⁢(𝒑)−ℒ︀⁢(𝒒,𝐠−1⁢(𝒑),t))\displaystyle{}=\frac{\mathrm{d}{}}{\mathrm{d}{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}}^{\text{T}}}% \mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}^{\text{T}}{\mathbf{g}}^{-1}({\color[rgb% ]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}})-% \mathcal{L}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{q}},{\mathbf{g}}^{-1}({\color[rgb]{.75,0,.25}\definecolor[named]% {pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}),t)}\right)
Eq. A.7
=𝐠−1⁢(𝒑)T+𝒑T⁢∂𝐠−1∂𝒑T⁢(𝒑)−∂ℒ︀∂𝒒˙T⁢∂𝐠−1∂𝒑T\displaystyle{}={\mathbf{g}}^{-1}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}})^{\text{T}}+{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}^{\text{T}}\frac{% \partial{{\mathbf{g}}^{-1}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}}}^{\text{T}}}({\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}})-\frac{% \partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}}}}^{\text{T}}}\frac{\partial{{% \mathbf{g}}^{-1}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}}}^{\text{T}}}
=𝐠−1⁢(𝒑)T\displaystyle{}={\mathbf{g}}^{-1}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}})^{\text{T}}
Eq. A.3.

In short, we have

equation (A.8) (A.8)
𝒒˙=∂ℋ︀∂𝒑⁢(𝒒,𝒑,t),\displaystyle{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\dot{q}}}=\frac{\partial{\mathcal{H}}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}}}}({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{q% }},{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{p}},t),
−𝒑˙∗=∂ℋ︀∂𝒒⁢(𝒒∗,𝒑∗,t).\displaystyle-\dot{\bm{p}}^{*}=\frac{\partial{\mathcal{H}}}{\partial{{\color[% rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]% {.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}}}}}(\bm{q% }^{*},\bm{p}^{*},t).

As with Eq. A.3, the first equation holds for all trajectories, whereas the second holds only for optimal trajectories 𝒒∗⁢(t)\bm{q}^{*}(t). Indeed, the first equation gives us the inverse function 𝐠−1{\mathbf{g}}^{-1} mapping from momenta back to positions. That is, ∂ℋ︀/∂𝒑\partial{\mathcal{H}}/\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}}} and ∂ℒ︀/∂𝒒˙\partial{\mathcal{L}}/\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{q}}}}} are inverses in their second arguments.

Time evolution of the optimal Hamiltonian.

The relationships expressed by Eq. A.8 tell us something interesting about the total derivative of the Hamiltonian with respect to time:

dd⁢t⁢ℋ︀⁢(𝒒∗⁢(t),𝒑∗⁢(t),t)\displaystyle\frac{\mathrm{d}{}}{\mathrm{d}{t}}\mathcal{H}(\bm{q}^{*}(t),\bm{p% }^{*}(t),t)
=∂ℋ︀∂𝒒T⁢(𝒒∗,𝒑∗,t)⁢𝒒˙∗+∂ℋ︀∂𝒑T⁢(𝒒∗,𝒑∗,t)⁢𝒑˙∗+∂ℋ︀∂t⁢(𝒒∗,𝒑∗,t)\displaystyle{}=\frac{\partial{\mathcal{H}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}}}}^{\text{T}}}(\bm{q% }^{*},\bm{p}^{*},t)\bm{\dot{q}}^{*}+\frac{\partial{\mathcal{H}}}{\partial{{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p% }}}}^{\text{T}}}(\bm{q}^{*},\bm{p}^{*},t)\dot{\bm{p}}^{*}+\frac{\partial{% \mathcal{H}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}t}}}(\bm{q}^{*},\bm{p}^{*},t)
=−𝒑˙∗T⁢𝒒˙∗+𝒒˙∗T⁢𝒑˙∗+∂ℋ︀∂t⁢(𝒒∗,𝒑∗,t)\displaystyle{}=-{\dot{\bm{p}}^{*}{}}^{\text{T}}\bm{\dot{q}}^{*}+{\bm{\dot{q}}% ^{*}{}}^{\text{T}}\dot{\bm{p}}^{*}+\frac{\partial{\mathcal{H}}}{\partial{{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}t}}}(% \bm{q}^{*},\bm{p}^{*},t)
Eq. A.8
=∂ℋ︀∂t⁢(𝒒∗,𝒑∗,t).\displaystyle{}=\frac{\partial{\mathcal{H}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}t}}}(\bm{q}^{*},\bm{p}^{*},% t).

That is, the total and partial derivatives of the Hamiltonian coincide (for optimal trajectories); or, put another way, Hamiltonians that do not depend explicitly on time (autonomous Hamiltonians) are constant along their optimal trajectories.

Our expression for the Hamiltonian in terms of the Lagrangian (Eq. A.7) allows us to relate their time derivatives:

∂ℋ︀∂t⁢(𝒒,𝒑,t)=dd⁢t⁢(𝒑T⁢𝐠−1⁢(𝒑)−ℒ︀⁢(𝒒,𝐠−1⁢(𝒑),t))=−∂ℒ︀∂t⁢(𝒒,𝐠−1⁢(𝒑),t).\frac{\partial{\mathcal{H}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}t}}}({\color[rgb]{.75,0,.25}\definecolor% [named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}},{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}},t)=\frac{\mathrm{d}% {}}{\mathrm{d}{t}}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}^{\text{T}}{\mathbf{% g}}^{-1}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{p}})-\mathcal{L}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}},{\mathbf{g}}^{-1}({\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}),t)}% \right)=-\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}t}}}({\color[rgb]{.75,0,.25% }\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{q}},{\mathbf{g}}^{-1}(% {\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% p}}),t).

But putting this together with the previous equation also tells us that

dd⁢t⁢ℋ︀⁢(𝒒∗⁢(t),𝒑∗⁢(t),t)=−∂ℒ︀∂t⁢(𝒒∗⁢(t),𝐠−1⁢(𝒑∗⁢(t)),t).\frac{\mathrm{d}{}}{\mathrm{d}{t}}\mathcal{H}(\bm{q}^{*}(t),\bm{p}^{*}(t),t)=-% \frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}t}}}(\bm{q}^{*}(t),{\mathbf{g}}^{-1}(\bm% {p}^{*}(t)),t).

In fine, the (optimal) Lagrangian can change with time due to changes in the positions and momenta, as well as independently of those variables. But the changes due to position and momentum changes do not transmit into the Hamiltonian; they cancel each other. This is intuitive under our interpretations of the Lagrangian and Hamiltonian as, respectively, the difference and sum of the kinetic and potential energies: Changes in position and momentum shift energy between kinetic and potential, but do not change the total energy. But such changes obviously do alter the difference between kinetic and potential energies.

Why conservative dynamics.

We saw above that our Lagrangian, Eq. A.5, gives rise to Newton’s laws, but without any velocity-dependent force, i.e. without energy dissipation. In fact, no scalar potential function can generate such forces. To see this, let us apply the ELEs, Eq. A.1, to a generic Lagrangian. For brevity, we employ our shorthand for the derivative of the Lagrangian with respect to velocity (Eq. A.6), and omit arguments:

0=∂ℒ︀∂𝒙−(∂𝐠∂𝒙⁢𝒙˙∗+∂𝐠∂𝒙˙⁢𝒙¨∗+∂𝐠∂t)0=\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}}}-\mathopen{}\mathclose{{}\left% (\frac{\partial{\mathbf{g}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}}\bm{\dot{x}}^{*}+\frac{\partial% {\mathbf{g}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}\ddot{\bm{x}}^{*}+\frac{\partial% {\mathbf{g}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}t}}}}\right)

….††margin: FINISH THIS SECTION

A.2.2 The method of adjoints

Consider a possibly nonlinear dynamical system, 𝒙˙=𝐟⁢(𝒙,t),{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \dot{x}}}=\mathbf{f}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor% }{rgb}{.75,0,.25}\bm{x}},t), that accrues a running cost at the rate ℓ⁢(𝒙⁢(t),t)\ell({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25% }\bm{x}}(t),t). Let us define the (augmented) Lagrangian to be the sum of this running cost with Lagrange multipliers enforcing the constraint that the system obey the dynamics:

ℒ︀(𝒙,𝒙˙,t) . . =ℓ(𝒙,t)+𝝀T(𝒙˙−𝐟(𝒙,t)),\mathcal{L}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}},{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{% rgb}{.75,0,.25}\bm{\dot{x}}},t)\mathrel{\vbox{\hbox{.}\hbox{.} }}=\ell({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}},t)+\bm{\lambda}^{\text{T}}\mathopen{}\mathclose{{}\left({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \dot{x}}}-\mathbf{f}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor% }{rgb}{.75,0,.25}\bm{x}},t)}\right),

with 𝝀\bm{\lambda} the Lagrange multipliers. Notice that this setup differs conceptually from that of the mechanical systems in the previous section, in which the running cost, rather than the constraint, enforces the system dynamics. However, this poses no mathematical obstacle: Integrating this Lagrangian over time gives the total cost, which is a functional of the function 𝒙⁢(t)\bm{x}(t). Therefore the optimal trajectory is obtainable with the calculus of variations.

We will reformulate this in terms of a Hamiltonian. Applying the definition of the conjugate variables (Eq. A.3) to this Lagrangian, we obtain

𝒑 . . =∂ℒ︀∂𝒙˙(𝒙,𝒙˙,t)=𝝀.{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% p}}\mathrel{\vbox{\hbox{.}\hbox{.} }}=\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}},{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}},t)=% \bm{\lambda}.

That is, the conjugate variables are the Lagrange multipliers. In this context, they are known as the costate††margin: costate . Next, we form the Hamiltonian according to the recipe in Eq. A.7:

equation (A.9) (A.9)
ℋ︀⁢(𝒙,𝒑,t)=𝒑T⁢𝐠−1⁢(𝒑)−ℒ︀⁢(𝒙,𝐠−1⁢(𝒑),t)=𝒑T⁢𝐟⁢(𝒙,t)−ℓ⁢(𝒙,t).\mathcal{H}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}},{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{% rgb}{.75,0,.25}\bm{p}},t)={\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}^{\text{T}}{\mathbf{g}}^{-1}({\color[rgb% ]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}})-% \mathcal{L}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}},{\mathbf{g}}^{-1}({\color[rgb]{.75,0,.25}\definecolor[named]% {pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}),t)={\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{p}}^{\text{T}}\mathbf{f% }({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{x}},t)-\ell({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb% }{.75,0,.25}\bm{x}},t).

Finally, the optimal trajectories of the states 𝒙∗⁢(t)\bm{x}^{*}(t) and costates 𝒑∗⁢(t)\bm{p}^{*}(t) must obey the ELEs—in terms of the Hamiltonian, Eq. A.8. Applying it to Eq. A.9, we obtain

equation (A.10) (A.10)
𝒙˙=𝐟⁢(𝒙,t);\displaystyle{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\dot{x}}}=\mathbf{f}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}},t);
𝒑˙∗=∂ℓ∂𝒙⁢(𝒙∗,t)−𝒑∗⋅∂𝐟∂𝒙⁢(𝒙∗,t).\displaystyle\dot{\bm{p}}^{*}=\frac{\partial{\ell}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}}(\bm{x}^% {*},t)-\bm{p}^{*}\cdot\frac{\partial{\mathbf{f}}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}}(\bm{x}^% {*},t).

The first equation tells us nothing new, but the second tell us how the costate evolves with time. Typically, we are given an initial state (which we can take to be 𝟎\bm{0} w.l.o.g.), in which case the state dynamics must be integrated forward in time. If we have a terminal cost, then the transversality condition from Eq. A.3 provides the final costate, and therefore the costate dynamics must be integrated backwards in time.††margin: A word on numerical solution?

A.2.3 Pontryagin’s minimum principle

[[[In fact, if you insist on using the sign convention for the Hamiltonian from classical mechanics, as you have below, you get the maximum principle: we need to maximize, rather than minimize, the Hamiltonian.]]] We can generalize the foregoing construction along different avenues. One possibility is to let the dynamics—and the running cost—depend on a time-varying control input, 𝒖⁢(t)\bm{u}(t):

ℒ︀(𝒙,𝒙˙,𝒖,t) . . =ℓ(𝒙,𝒖,t)+𝝀T(𝒙˙−𝐟(𝒙,𝒖,t)).\mathcal{L}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}},{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{% rgb}{.75,0,.25}\bm{\dot{x}}},{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{u}},t)\mathrel{\vbox{\hbox{.}\hbox{.} }}=\ell({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}},{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{% rgb}{.75,0,.25}\bm{u}},t)+\bm{\lambda}^{\text{T}}\mathopen{}\mathclose{{}\left% ({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm% {\dot{x}}}-\mathbf{f}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}},{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{u}},t)}\right).

This is an optimal control problem: we seek the controls that steer the system through some desirable trajectory, also perhaps subject to some constraint on the controls (e.g., that they not be too large). The development is essentially identical to that of the method of adjoints just given. The only difference is that we derive an additional PDE, for the controls. It can be treated just like the state and in general could generate a matching PDE. However, we typically do not expect the cost or dynamics to depend directly on the time derivative of the controls, 𝒖˙\dot{\bm{u}}, so the Euler-Lagrange equations (Eq. A.1) for the controls reduce to

equation (A.11) (A.11)
0=∂ℒ︀∂𝒖⁢(𝒙∗⁢(t),𝒙˙∗⁢(t),𝒖∗⁢(t),t)=∂ℋ︀∂𝒖⁢(𝒙∗⁢(t),𝒑∗⁢(t),𝒖∗⁢(t),t).0=\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{u}}}}(\bm{x}^{*}({\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}t}),\bm{\dot{x}}^% {*}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% t}),\bm{u}^{*}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}% {.75,0,.25}t}),{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}% {.75,0,.25}t})=\frac{\partial{\mathcal{H}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{u}}}}(\bm{x}^{*}({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}t}),% \bm{p}^{*}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}t}),\bm{u}^{*}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}t}),{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}t}).

The fact the Lagrangian and Hamiltonian gradients are equal at the optimal control follows from the recipe for constructing Hamiltonians, Eq. A.7.

How do we actually use these equations to solve for optimal controls? The entire optimal system is now governed by three PDEs: Eq. A.10 (modified to include the controls) and Eq. A.11. ….††margin: FINISH THIS SECTION

A.2.4 Neural ODEs

Another possible avenue for generalization is to consider optimizing for fixed (time-invariant) parameters 𝜽\bm{\theta} (rather than time-varying controls). In this framework, the “neural ODEs” of Chen and colleagues [6] can be33 3 Those authors took a slightly different approach. derived. The authors consider a certain continuous-time limit of recurrent neural networks (or continuous-layer limit of residual networks), in which forward propagation becomes a nonlinear dynamical system,

𝒙t+1=𝒙t+𝐟⁢(𝒙t,𝜽t)⟶𝒙˙⁢(t)=𝐟⁢(𝒙⁢(t),𝜽,t).\bm{x}_{t+1}=\bm{x}_{t}+\mathbf{f}(\bm{x}_{t},\bm{\theta}_{t})\longrightarrow% \bm{\dot{x}}(t)=\mathbf{f}(\bm{x}(t),\bm{\theta},t).

What does backpropagation become?

We can use the same Lagrangian as in the method of adjoints above, so the state and costate dynamics are again given by Eq. A.10. The terminal cost 𝒮︀\mathcal{S} can be identified with the loss function, applied at the final layer/time, although we can impose costs at other layers/times through the running cost ℓ\ell. We also have in hand boundary values: the initial state is the input to the network (or some transformation thereof), and the final costate is the loss gradient evaluated at the network outputs (from the transversality condition, Eq. A.2):

𝒙⁢(0)=𝒚,\displaystyle\bm{x}(0)=\bm{y},
𝒑∗⁢(T)=−∂𝒮︀∂𝒙⁢(𝒙∗⁢(T)).\displaystyle\bm{p}^{*}(T)=-\frac{\partial{\mathcal{S}}}{\partial{{\color[rgb]% {.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}}(\bm{x}% ^{*}(T)).

Therefore it is possible to compute, or at least numerically simulate, the trajectories of the state and of the costate, by first a forward pass through the state dynamics, followed by a backward pass with the costate (adjoint) dynamics. Comparing Eq. A.10 to Eqs. 7.19 and 7.21, we see that we have derived precisely the continuous-time analogs of the forward and backward passes required for backpropagation through time.

To improve the model, we need the gradient d⁢J/d⁢𝜽\mathrm{d}{J}/\mathrm{d}{\bm{\theta}} of the (total) loss. Here we make use of an envelope theorem (cf. Eq. A.4). Along optimal trajectories of the states 𝒙\bm{x} and their costates 𝒑\bm{p},††margin: We leave the extension of the envelope theorem to the costate as an exercise for the reader: Exercise LABEL:ex:We_leave_the_extension_of_the_envelope_theorem_to_the_costate_as_an_exercise_for_the_reader the loss function’s total derivative is just its partial derivative:

equation (A.12) (A.12)
dd⁢𝜽⁢J⁢[𝒙∗]=∂𝒮︀∂𝜽⁢(𝒙∗⁢(T),𝜽)+∫0T∂ℒ︀∂𝜽⁢(𝒙∗⁢(t),𝒙˙∗⁢(t),𝜽,t)⁢dt=∂𝒮︀∂𝜽⁢(𝒙∗⁢(T),𝜽)−∫0T𝒑⁢(t)⋅∂𝐟∂𝜽⁢(𝒙⁢(t),𝜽,t)⁢dt.\begin{split}\frac{\mathrm{d}{}}{\mathrm{d}{\bm{\theta}}}J[\bm{x}^{*}]&{}=% \frac{\partial{\mathcal{S}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}}}}(\bm{x}^{*}(T),{\color[rgb% ]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}})+% \int_{0}^{T}\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}}}}(\bm{x}^{*}(t% ),\bm{\dot{x}}^{*}(t),{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}},t)\mathop{}\!\mathrm{d}{t{}}\\ &{}=\frac{\partial{\mathcal{S}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}}}}(\bm{x}^{*}(T),{\color[rgb% ]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}})-% \int_{0}^{T}\bm{p}(t)\cdot\frac{\partial{\mathbf{f}}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}}}}(% \bm{x}(t),{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\theta}},t)\mathop{}\!\mathrm{d}{t{}}.\end{split}

This equation provides a simple formula for optimization of the parameters. Indeed, it can be numerically integrated, also in the backwards direction, along with 𝒑\bm{p}.

Notice that there is no need to backpropagate gradients through the ODE solver. The authors also suggest re-solving for the states 𝒙\bm{x} in the backward pass to avoid having to store them during the forward pass. (A forward pass is necessary in any case in order to initialize the final state.)

A.2.5 Hamiltonian Monte Carlo

††margin: FINISH THIS SECTION

A.2.6 Exponential families

Consider an exponential family of probability distributions,

p^⁢(𝒙;𝜼)=h⁢(𝒙)⁢exp⁡{𝜼⋅𝑻⁢(𝒙)−A⁢(𝜼)}.{\hat{p}\mathopen{}\mathclose{{}\left({\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}{};{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\eta}}}\right)}=h({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x% }})\exp\mathopen{}\mathclose{{}\left\{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\eta}}\cdot{\bm{T}}{}({\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}})-A({% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \eta}})}\right\}.

It is well known (and easily shown) that the first and second derivatives of the log-partition function A⁢(𝜼)A({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{\eta}}) are (respectively) the expectation and covariance of the sufficient statistics:

∂A∂𝜼=𝔼p^[𝑻(𝑿^)]= . . 𝝁,\displaystyle\frac{\partial{A}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\eta}}}}}=\mathbb{E}_{\hat{p}}{% \mathopen{}\mathclose{{}\left[{\bm{T}}{}({\bm{\hat{X}}})}\right]}=\mathrel{% \vbox{\hbox{.}\hbox{.} }}{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{\mu}},
∂2A∂𝜼⁢∂𝜼T=Covp^⁢[𝑻⁢(𝑿^)].\displaystyle\frac{\partial^{2}{A}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\eta}}}}\partial{{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \eta}}}}^{\text{T}}}=\text{Cov}_{\hat{p}}{\mathopen{}\mathclose{{}\left[{\bm{T% }}{}({\bm{\hat{X}}})}\right]}.

Since covariance matrices are positive definite††margin: What about degenerate cases with only positive semi-definite? , the first derivative provides an invertible map between the natural parameters and the moment parameters:

equation (A.13) (A.13)
𝝁=∂A∂𝜼(𝜼)= . . 𝐠(𝜼);{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \mu}}=\frac{\partial{A}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\eta}}}}}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\eta}})=\mathrel{\vbox{% \hbox{.}\hbox{.} }}\mathbf{g}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\eta}});

and the log-partition function AA is itself convex in the natural parameters.

This suggests a certain analogy between the Lagrangian and the log-partition function, with the natural parameters playing the role of velocities and the moment parameters like momenta. Let us take the analogy seriously and carry out a Legendre transformation, exactly on analogy with Eq. A.7

A⁢(𝜼)\displaystyle A({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb% }{.75,0,.25}\bm{\eta}})
=∂A∂𝜼T⁢𝜼−ℋ︀⁢(𝐠⁢(𝜼))\displaystyle{}=\frac{\partial{A}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\eta}}}}^{\text{T}}}{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \eta}}-\mathcal{H}(\mathbf{g}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\eta}}))
y=m⁢x+b\displaystyle{}y=mx+b
⟹ℋ︀⁢(𝝁)\displaystyle\implies\mathcal{H}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\mu}})
=𝝁T⁢𝐠−1⁢(𝝁)−A⁢(𝐠−1⁢(𝝁))\displaystyle{}={\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb% }{.75,0,.25}\bm{\mu}}^{\text{T}}{\mathbf{g}}^{-1}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\mu}})-A({\mathbf{g}}^{% -1}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{\mu}}))
Eq. A.13.

Notice that, as in Eq. A.8,

equation (A.14) (A.14)
∂ℋ︀∂𝝁⁢(𝝁)=𝜼.\frac{\partial{\mathcal{H}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\mu}}}}}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\mu}})={\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\eta}}.

However, we can also begin from a different point. Let us simply define the “Hamiltonian” to be the relative entropy of the distribution and its base measure:44 4 The base measure need not integrate to one, i.e. be a proper probability density, so the minimum of the Hamiltonian need not be 0.

equation (A.15) (A.15)
ℋ︀⁢(𝝁) . . =DKL{p^(𝑿^;𝜽)∥h(𝑿^)}=𝔼𝑿^⁢[log⁡h⁢(𝑿^)+𝐠−1⁢(𝝁)⋅𝑻⁢(𝑿^)−A⁢(𝐠−1⁢(𝝁))−log⁡h⁢(𝑿^)]=𝐠−1⁢(𝝁)⋅𝝁−A⁢(𝐠−1⁢(𝝁)).\begin{split}\mathcal{H}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\mu}})&{}\mathrel{\vbox{\hbox{.}\hbox{.} }}=\operatorname*{\text{D}_{\text{KL}}}\mathopen{}\mathclose{{}\left\{{\hat{p}% \mathopen{}\mathclose{{}\left({\bm{\hat{X}}};\bm{\theta}}\right)}\middle\|{h% \mathopen{}\mathclose{{}\left({\bm{\hat{X}}}}\right)}}\right\}\\ &{}=\mathbb{E}_{{\bm{\hat{X}}}{}}{\mathopen{}\mathclose{{}\left[\log h({\bm{% \hat{X}}})+{\mathbf{g}}^{-1}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\mu}})\cdot{\bm{T}}{}({\bm{\hat{X}}})-A({% \mathbf{g}}^{-1}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{% rgb}{.75,0,.25}\bm{\mu}}))-\log h({\bm{\hat{X}}})}\right]}\\ &{}={\mathbf{g}}^{-1}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\mu}})\cdot{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\mu}}-A({\mathbf{g}}^{-% 1}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{\mu}})).\end{split}

This yields the same result as the Legendre transformation, and consequently Eqs. A.13 and A.14 still hold.