Path-Integral Value-Control Matching with Applications to Smoothing and Game
Abstract
Stochastic optimal control (SOC) provides a computational framework for a broad range of machine-learning problems. Policy-based solvers learn feedback controls directly, but often require costly full-horizon rollouts and may suffer from unstable likelihood-ratio weights during off-policy training. Local path-integral value learning enables short-trajectory updates and experience replay, but its scalar objective does not directly supervise the value gradient that determines the control. This limitation is particularly relevant to control-sensitive applications such as Bayesian smoothing and mean-field games. Building on Path-Integral Value Matching (PI-VM), we propose Path-Integral Value-Control Matching (PI-VCM), which adds direct control supervision while retaining local value learning and replay. We spatially differentiate the local path-integral value recursion to derive a corresponding control recursion, allowing the same short trajectories to provide both value and control targets for a shared value network. Affine sensitivity caching and local likelihood correction further enable off-policy experience replay without repeated adjoint propagation. Experiments demonstrate improved control accuracy over baselines on SOC benchmarks and higher-dimensional instances, together with competitive performance on Bayesian smoothing and mean-field games.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.