A.1 The Calculus of Variations

We first encounter optimization problems in introductory calculus, in which we seek the extreme value of some function. The standard method is to find the point(s) at which the gradient is zero—a necessary condition for optimality. If the point we seek is not free but rather confined to a line or surface, then the problem becomes slightly more complicated: Noting that a constraint surface f⁢(𝒙)=cf(\bm{x})=c and its gradient are perpendicular, and that the gradient of a function points “uphill,” we conclude that the optimal point on the constraint surface must occur where these two gradients are aligned, since at this point movement along the surface will not go up (or down) hill. This can be formalized as the method of Lagrange multipliers.

A more complicated problem still arises when we seek an optimal trajectory, 𝒙⁢(t)\bm{x}(t). Intuitively, the function we wish to extremize can be thought of as a cost that accumulates over the course of the entire trajectory,

J⁢[𝒙]=𝒮︀⁢(𝒙⁢(T))+∫0Tℒ︀⁢(𝒙⁢(t),𝒙˙⁢(t),t)⁢dt,J[{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{x}}]=\mathcal{S}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor% }{rgb}{.75,0,.25}\bm{x}}(T))+\int_{0}^{T}\mathcal{L}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}(t),{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}(t),% t)\mathop{}\!\mathrm{d}{t{}},

with the Lagrangian ℒ︀\mathcal{L} providing the instantaneous or running cost rate. In this formulation we have also allowed for a terminal cost††margin: terminal cost , 𝒮︀⁢(𝒙⁢(T))\mathcal{S}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}(T)), including any constraints. We will also assume a known initial condition, 𝒙⁢(0)\bm{x}(0), which we let be (w.l.o.g.) 𝟎\bm{0}.

The Euler-Lagrange equations and the transversality condition.

Since our object is now the function that optimizes the functional JJ, rather than the point that optimizes the function ℒ︀\mathcal{L}, we need a more sophisticated method. The basic approach is to transform this optimization into the simple one over scalars from introductory calculus. In particular, consider small displacements to the optimal trajectory 𝒙∗\bm{x}^{*} by another (nicely behaved) trajectory 𝜼⁢(t)\bm{\eta}(t); i.e. consider trajectories 𝒙∗⁢(t)+ϵ⁢𝜼⁢(t)\bm{x}^{*}(t)+\epsilon\bm{\eta}(t). Thus, we have expressed a family of trajectories near the optimum in terms of a scalar parameter ϵ\epsilon, which allows us to optimize JJ with the standard method of single-variable calculus. The additional wrinkle is that the optimization must hold for any choice of displacement trajectories 𝜼⁢(t)\bm{\eta}(t). However, we do still require that the displaced trajectory satisfy the boundary conditions, in this case that 𝒙∗⁢(0)+ϵ⁢𝜼⁢(0)=𝟎\bm{x}^{*}(0)+\epsilon\bm{\eta}(0)=\bm{0}, which implies that 𝜼⁢(0)=0\bm{\eta}(0)=0.††margin: Precisely why is this necessary?

Accordingly, we set the derivative equal to zero and then proceed with some lemmas from multivariable calculus:

0\displaystyle 0
=setdd⁢ϵJ[𝒙∗+ϵ𝜼(t)]|ϵ=0\displaystyle{}\stackrel{{\scriptstyle\text{set}}}{{=}}\frac{\mathrm{d}{}}{% \mathrm{d}{\epsilon}}\mathopen{}\mathclose{{}\left.J[\bm{x}^{*}+\epsilon\bm{% \eta}(t)]}\right\rvert_{\epsilon=0}
=dd⁢ϵ(𝒮︀(𝒙∗(T)+ϵ𝜼(T))+∫0Tℒ︀(𝒙∗(t)+ϵ𝜼(t),𝒙˙∗(t)+ϵ𝜼˙(t),t)dt)|ϵ=0\displaystyle{}=\mathopen{}\mathclose{{}\left.\frac{\mathrm{d}{}}{\mathrm{d}{% \epsilon}}\mathopen{}\mathclose{{}\left(\mathcal{S}(\bm{x}^{*}(T)+\epsilon\bm{% \eta}(T))+\int_{0}^{T}\mathcal{L}(\bm{x}^{*}(t)+\epsilon\bm{\eta}(t),\bm{\dot{% x}}^{*}(t)+\epsilon\dot{\bm{\eta}}(t),t)\mathop{}\!\mathrm{d}{t{}}}\right)}% \right\rvert_{\epsilon=0}
=∂𝒮︀∂𝒙T⁢𝜼⁢(T)+∫0T(∂ℒ︀∂𝒙T⁢𝜼⁢(t)+∂ℒ︀∂𝒙˙T⁢𝜼˙⁢(t))⁢dt\displaystyle{}=\frac{\partial{\mathcal{S}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}^{\text{T}}}\bm{% \eta}(T)+\int_{0}^{T}\mathopen{}\mathclose{{}\left(\frac{\partial{\mathcal{L}}% }{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}}^{\text{T}}}\bm{\eta}(t)+\frac{\partial{\mathcal{L}}}{% \partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\dot{x}}}}^{\text{T}}}\dot{\bm{\eta}}(t)}\right)\mathop{}\!% \mathrm{d}{t{}}
Leibniz’s rule
=∂𝒮︀∂𝒙T𝜼(T)+∂ℒ︀∂𝒙˙T𝜼(t)|t=0t=T+∫0T(∂ℒ︀∂𝒙T𝜼(t)−(dd⁢t∂ℒ︀∂𝒙˙T)𝜼(t))dt\displaystyle{}=\frac{\partial{\mathcal{S}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}^{\text{T}}}\bm{% \eta}(T)+\mathopen{}\mathclose{{}\left.\frac{\partial{\mathcal{L}}}{\partial{{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \dot{x}}}}^{\text{T}}}\bm{\eta}(t)}\right\rvert_{t=0}^{t=T}+\int_{0}^{T}% \mathopen{}\mathclose{{}\left(\frac{\partial{\mathcal{L}}}{\partial{{\color[% rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}^{% \text{T}}}\bm{\eta}(t)-\mathopen{}\mathclose{{}\left(\frac{\mathrm{d}{}}{% \mathrm{d}{t}}\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}^{\text{T}}}}% \right)\bm{\eta}(t)}\right)\mathop{}\!\mathrm{d}{t{}}
integ. by parts
=(∂𝒮︀∂𝒙+∂ℒ︀∂𝒙˙)T⁢𝜼⁢(T)−∂ℒ︀∂𝒙˙T⁢𝜼⁢(0)+∫0T(∂ℒ︀∂𝒙−dd⁢t⁢∂ℒ︀∂𝒙˙)T⁢𝜼⁢(t)⁢dt.\displaystyle{}=\mathopen{}\mathclose{{}\left(\frac{\partial{\mathcal{S}}}{% \partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{x}}}}+\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}}% \right)^{\text{T}}\bm{\eta}(T)-\frac{\partial{\mathcal{L}}}{\partial{{\color[% rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}% }}^{\text{T}}}\bm{\eta}(0)+\int_{0}^{T}\mathopen{}\mathclose{{}\left(\frac{% \partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}}-\frac{\mathrm{d}{}}{\mathrm{d}{t}}% \frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}}\right)^{\text{T}}\bm{% \eta}(t)\mathop{}\!\mathrm{d}{t{}}.

The second term vanishes under the assumption that the displaced trajectory satisfy the boundary condition. The remaining terms involve inner products with the displacement trajectory 𝜼⁢(t)\bm{\eta}(t). Since the equation must hold for any choice of displacements trajectories, the other vectors in these inner products must be zero. Making this explicit (and restoring the suppressed arguments of functions), we obtain the Euler-Lagrange equations††margin: Euler-Lagrange equations ,

equation (A.1) (A.1)
∂ℒ︀∂𝒙⁢(𝒙∗⁢(t),𝒙˙∗⁢(t),t)−dd⁢t⁢∂ℒ︀∂𝒙˙⁢(𝒙∗⁢(t),𝒙˙∗⁢(t),t)=0,\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}}}(\bm{x}^{*}({\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}t}),\bm{\dot{x}}^% {*}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% t}),{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% t})-\frac{\mathrm{d}{}}{\mathrm{d}{t}}\frac{\partial{\mathcal{L}}}{\partial{{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \dot{x}}}}}}(\bm{x}^{*}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}t}),\bm{\dot{x}}^{*}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}t}),{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}t})=0,

and the transversality condition††margin: transversality condition ,

equation (A.2) (A.2)
∂𝒮︀∂𝒙⁢(𝒙∗⁢(T))+∂ℒ︀∂𝒙˙⁢(𝒙∗⁢(T),𝒙˙∗⁢(T),T)=0.\frac{\partial{\mathcal{S}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}}(\bm{x}^{*}(T))+\frac{\partial{% \mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}}(\bm{x}^{*}(T),\bm{\dot{x}}^{*}% (T),T)=0.

In use of the ELEs, the partial derivative ∂ℒ︀/∂𝒙˙\partial{\mathcal{L}}/\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}} will play a special role. Therefore we endow it with a name, the conjugate variable††margin: conjugate variable 𝒑\bm{p}, and re-write Eqs. A.1 and A.2 in terms of it:

equation (A.3) (A.3)
𝒑(t) . . =∂ℒ︀∂𝒙˙(𝒙(t),𝒙˙(t),t),\displaystyle\bm{p}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}% {rgb}{.75,0,.25}t})\mathrel{\vbox{\hbox{.}\hbox{.} }}=\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}(\bm{x}({\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}t}),\bm{\dot{x}}(% {\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}t}),% {\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}t}),
𝒑˙⁢(t)=∂ℒ︀∂𝒙⁢(𝒙∗⁢(t),𝒙˙∗⁢(t),t)\displaystyle\dot{\bm{p}}({\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}t})=\frac{\partial{\mathcal{L}}}{\partial{{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x% }}}}}(\bm{x}^{*}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{% rgb}{.75,0,.25}t}),\bm{\dot{x}}^{*}({\color[rgb]{.75,0,.25}\definecolor[named]% {pgfstrokecolor}{rgb}{.75,0,.25}t}),{\color[rgb]{.75,0,.25}\definecolor[named]% {pgfstrokecolor}{rgb}{.75,0,.25}t})
𝒑⁢(T)=−∂𝒮︀∂𝒙⁢(𝒙∗⁢(T)).\displaystyle\bm{p}(T)=-\frac{\partial{\mathcal{S}}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}}(\bm{x}^% {*}(T)).

The first two equations have a nice symmetry. Note well that the first is a definition and holds for all trajectories 𝒙⁢(t){\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% x}}(t), whereas the second and third hold only for optimal trajectories 𝒙∗⁢(t)\bm{x}^{*}(t).

The envelope theorem.

Let the Lagrangian and terminal cost be parameterized with some time-invariant parameters 𝜽\bm{\theta}, which we wish to optimize. Then we require the total derivative of the accumlated loss JJ,

dd⁢𝜽⁢J⁢[𝒙]=dd⁢𝜽⁢(𝒮︀⁢(𝒙⁢(T),𝜽)+∫0Tℒ︀⁢(𝒙⁢(t),𝒙˙⁢(t),𝜽,t)⁢dt)=(∂𝒮︀∂𝒙d⁢𝒙d⁢𝜽+∂𝒮︀∂𝜽)|t=T+∫0T(∂ℒ︀∂𝒙d⁢𝒙d⁢𝜽+∂ℒ︀∂𝒙˙d⁢𝒙˙d⁢𝜽+∂ℒ︀∂𝜽)dt.\begin{split}\frac{\mathrm{d}{}}{\mathrm{d}{\bm{\theta}}}J[{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}]&{}=\frac% {\mathrm{d}{}}{\mathrm{d}{\bm{\theta}}}\mathopen{}\mathclose{{}\left(\mathcal{% S}({\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{x}}(T),{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\theta}})+\int_{0}^{T}\mathcal{L}({\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}(t),{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}(t),% {\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \theta}},t)\mathop{}\!\mathrm{d}{t{}}}\right)\\ &{}=\mathopen{}\mathclose{{}\left.\mathopen{}\mathclose{{}\left(\frac{\partial% {\mathcal{S}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}}\frac{\mathrm{d}{\bm{x}}}{\mathrm{d}{% \bm{\theta}}}+\frac{\partial{\mathcal{S}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}}}}}\right)}% \right\rvert_{t=T}+\int_{0}^{T}\mathopen{}\mathclose{{}\left(\frac{\partial{% \mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{x}}}}\frac{\mathrm{d}{\bm{x}}}{\mathrm{d}{% \bm{\theta}}}+\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}\frac{% \mathrm{d}{\bm{\dot{x}}}}{\mathrm{d}{\bm{\theta}}}+\frac{\partial{\mathcal{L}}% }{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\theta}}}}}\right)\mathop{}\!\mathrm{d}{t{}}.\end{split}

(Arguments have been suppressed on the second line for brevity.) This is certainly complicated. However, at an optimal trajectory, 𝒙∗\bm{x}^{*}, the total derivative simplifies considerably. In particular, at an optimal trajectory, the Euler-Lagrange equations (Eq. A.1) and transversality condition (Eq. A.2) hold. Substituting these into the gradient, we find

equation (A.4) (A.4)
dd⁢𝜽⁢J⁢[𝒙∗]=dd⁢𝜽⁢𝒮︀⁢(𝒙∗⁢(T),𝜽)+∫0Tdd⁢𝜽⁢ℒ︀⁢(𝒙∗⁢(t),𝒙˙∗⁢(t),𝜽,t)⁢dt=(−∂ℒ︀∂𝒙˙d⁢𝒙d⁢𝜽+∂𝒮︀∂𝜽)|t=T+∫0T((dd⁢t∂ℒ︀∂𝒙˙)d⁢𝒙d⁢𝜽+∂ℒ︀∂𝒙˙d⁢𝒙˙d⁢𝜽+∂ℒ︀∂𝜽)dt=(−∂ℒ︀∂𝒙˙d⁢𝒙d⁢𝜽+∂𝒮︀∂𝜽)|t=T+(∂ℒ︀∂𝒙˙d⁢𝒙d⁢𝜽)|t=0t=T+∫0T(−∂ℒ︀∂𝒙˙d⁢𝒙˙d⁢𝜽+∂ℒ︀∂𝒙˙d⁢𝒙˙d⁢𝜽+∂ℒ︀∂𝜽)dt=∂𝒮︀∂𝜽|t=T−(∂ℒ︀∂𝒙˙d⁢𝒙d⁢𝜽)|t=0+∫0T∂ℒ︀∂𝜽dt=∂𝒮︀∂𝜽⁢(𝒙∗⁢(T))+∫0T∂ℒ︀∂𝜽⁢(𝒙∗,𝒙˙∗,𝜽,t)⁢dt.\begin{split}\frac{\mathrm{d}{}}{\mathrm{d}{\bm{\theta}}}J[\bm{x}^{*}]&{}=% \frac{\mathrm{d}{}}{\mathrm{d}{\bm{\theta}}}\mathcal{S}(\bm{x}^{*}(T),{\color[% rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}}% )+\int_{0}^{T}\frac{\mathrm{d}{}}{\mathrm{d}{\bm{\theta}}}\mathcal{L}(\bm{x}^{% *}(t),\bm{\dot{x}}^{*}(t),{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}},t)\mathop{}\!\mathrm{d}{t{}}\\ &{}=\mathopen{}\mathclose{{}\left.\mathopen{}\mathclose{{}\left(-\frac{% \partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}\frac{\mathrm{d}{\bm{x}}}{% \mathrm{d}{\bm{\theta}}}+\frac{\partial{\mathcal{S}}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}}}}}% \right)}\right\rvert_{t=T}+\int_{0}^{T}\mathopen{}\mathclose{{}\left(\mathopen% {}\mathclose{{}\left(\frac{\mathrm{d}{}}{\mathrm{d}{t}}\frac{\partial{\mathcal% {L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}% {.75,0,.25}\bm{\dot{x}}}}}}\right)\frac{\mathrm{d}{\bm{x}}}{\mathrm{d}{\bm{% \theta}}}+\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}\frac{% \mathrm{d}{\bm{\dot{x}}}}{\mathrm{d}{\bm{\theta}}}+\frac{\partial{\mathcal{L}}% }{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\theta}}}}}\right)\mathop{}\!\mathrm{d}{t{}}\\ &{}=\mathopen{}\mathclose{{}\left.\mathopen{}\mathclose{{}\left(-\frac{% \partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[named]{% pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}\frac{\mathrm{d}{\bm{x}}}{% \mathrm{d}{\bm{\theta}}}+\frac{\partial{\mathcal{S}}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}}}}}% \right)}\right\rvert_{t=T}+\mathopen{}\mathclose{{}\left.\mathopen{}\mathclose% {{}\left(\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}\frac{% \mathrm{d}{\bm{x}}}{\mathrm{d}{\bm{\theta}}}}\right)}\right\rvert_{t=0}^{t=T}+% \int_{0}^{T}\mathopen{}\mathclose{{}\left(-\frac{\partial{\mathcal{L}}}{% \partial{{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{% .75,0,.25}\bm{\dot{x}}}}}\frac{\mathrm{d}{\bm{\dot{x}}}}{\mathrm{d}{\bm{\theta% }}}+\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}\frac{\mathrm{d}{\bm{\dot% {x}}}}{\mathrm{d}{\bm{\theta}}}+\frac{\partial{\mathcal{L}}}{\partial{{\color[% rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}}% }}}\right)\mathop{}\!\mathrm{d}{t{}}\\ &{}=\mathopen{}\mathclose{{}\left.\frac{\partial{\mathcal{S}}}{\partial{{% \color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{% \theta}}}}}\right\rvert_{t=T}-\mathopen{}\mathclose{{}\left.\mathopen{}% \mathclose{{}\left(\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{% .75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\dot{x}}}}}% \frac{\mathrm{d}{\bm{x}}}{\mathrm{d}{\bm{\theta}}}}\right)}\right\rvert_{t=0}+% \int_{0}^{T}\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}% \definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}}}}\mathop{}\!% \mathrm{d}{t{}}\\ &{}=\frac{\partial{\mathcal{S}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}}}}(\bm{x}^{*}(T))+\int_{0}^{% T}\frac{\partial{\mathcal{L}}}{\partial{{\color[rgb]{.75,0,.25}\definecolor[% named]{pgfstrokecolor}{rgb}{.75,0,.25}\bm{\theta}}}}(\bm{x}^{*},\bm{\dot{x}}^{% *},{\color[rgb]{.75,0,.25}\definecolor[named]{pgfstrokecolor}{rgb}{.75,0,.25}% \bm{\theta}},t)\mathop{}\!\mathrm{d}{t{}}.\end{split}

To reach the third line, we integrated by parts. To reach the last line, we observed that the initial states 𝒙⁢(0)\bm{x}(0) are fixed, so d⁢𝒙/d⁢𝜽\mathrm{d}{\bm{x}}/\mathrm{d}{\bm{\theta}} must be 0. Crucially, the final line is the first line, with all total derivatives replaced by partial derivatives. Thus, along optimal trajectories, only the direct effects of parameters on terminal and running costs contribute to the total-cost gradient.