AdvG: A Provably Convergent Actor–Critic Algorithm for Reinforcement Learning in Continuous Time and Space
Abstract
Reinforcement learning in continuous time and state-action space poses a core challenge for today's RL research. Deterministic actor–critic methods offer a model-free approach by estimating how action changes affect expected return, while avoiding action-space integration in the policy update. For controlled diffusions with unknown dynamics, however, estimating these effects becomes difficult as shorter observation intervals reduce the signal-to-noise ratio. We propose AdvG, a novel deterministic policy gradient (DPG) algorithm using multi-step temporal difference error for value estimation while learning the advantage rate's action derivative from individual exploratory perturbations. By analyzing critic stability, we establish the first finite-sample joint actor and critic convergence guarantee in model-free DPG learning for nonlinear controlled diffusions. With exact critic features, we establish an joint convergence rate, with the expected total duration of collected RL trajectories. Furthermore, our algorithm inspires the construction of an actor gradient estimator with mean-squared error of order , achieving the minimax optimality over a family of diffusion models allowing nonlinear action effects.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.