Path Integral Value Matching: Regularized Control of Arbitrary Markov Processes
Abstract
KL-regularized control of continuous-time Markov processes seeks an optimal path distribution that balances expected path cost against deviation from a reference Markov process. Many existing solvers directly optimize control policies, but often require process-specific formulations and costly full-horizon rollouts, with off-policy reuse limited by accumulated likelihood-ratio variability. We propose Path Integral Value Matching (PI-VM), a process-agnostic value-based control method. Our method builds on a common path-integral representation of the value, which determines the optimal controlled dynamics through a Doob's -transform once the reference process is fixed. However, direct path-integral estimation still requires full-horizon trajectory simulation and can suffer from high variance. By truncating and marginalizing the path integral, we obtain a temporal recursion for the value function. PI-VM learns this recursion through temporal-difference learning on short trajectories, while experience replay and likelihood correction support off-policy training. The same learning objective applies to Euclidean and manifold diffusions, continuous-time Markov chains, and finite-activity jump processes. Experiments show that PI-VM achieves the best reported performance on multiple tasks across diverse process families and performs well in image fine-tuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.