acceptodds
Under review as a conference paper at ICLR 2027

Differentiating Through Delay: Stochastic Control Without Full History Hessians

Abstract

In stochastic delay control, the current state re-enters future dynamics and costs through the history. The stochastic maximum principle represents these effects through anticipated first-order adjoints and, with control-dependent diffusion, a coupled second-order family carrying current–history cross-curvature. We recover the information needed for constrained control refinement directly from a learned history policy's simulated continuations. Future control values are detached under state differentiation, while all physical delay paths remain active. Current-injection duality identifies the required diffusion curvature with an fixed-control Hessian block, eliminating full history-Hessian assembly and a separate numerical solution of the coupled adjoint family. The stored curvature is ; trajectory propagation retains its history dependence. We establish exact discrete adjoint correspondence and convergence of the population recovery inputs under the stated refinement assumptions. Building on Pontryagin-guided direct policy optimization (PGDPO), an anchored local solve then refines the control, with localized error bounds and sufficient conditions for objective improvement. Across five benchmarks, refinement reduces its learned reference policy's control error by 24–91 percent where optimizer references exist and significantly reduces paired cost on two constrained applications. Delay-model and mesh diagnostics test the computation separately from these performance comparisons.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.