Differentiating Through Delay: Stochastic Control Without Full History Hessians
Abstract
In stochastic delay control, the current state re-enters future dynamics and costs through the history. The stochastic maximum principle represents these effects through anticipated first-order adjoints and, with control-dependent diffusion, a coupled second-order family carrying current–history cross-curvature. We recover the information needed for constrained control refinement directly from a learned history policy's simulated continuations. Future control values are detached under state differentiation, while all physical delay paths remain active. Current-injection duality identifies the required diffusion curvature with an fixed-control Hessian block, eliminating full history-Hessian assembly and a separate numerical solution of the coupled adjoint family. The stored curvature is ; trajectory propagation retains its history dependence. We establish exact discrete adjoint correspondence and convergence of the population recovery inputs under the stated refinement assumptions. Building on Pontryagin-guided direct policy optimization (PGDPO), an anchored local solve then refines the control, with localized error bounds and sufficient conditions for objective improvement. Across five benchmarks, refinement reduces its learned reference policy's control error by 24–91 percent where optimizer references exist and significantly reduces paired cost on two constrained applications. Delay-model and mesh diagnostics test the computation separately from these performance comparisons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.