Diffusion-dependent regret bounds for continuous time reinforcement learnin
Abstract
Continuous-time reinforcement learning provides a natural framework for controlling dynamical systems, yet learning guarantees must account for both unknown dynamics and intrinsic stochasticity. We study continuous-time reinforcement learning in nonlinear stochastic systems with unknown drift and diffusion. The dynamics are governed by stochastic differential equations whose drift and diffusion terms admit linear parameterizations in known, potentially nonlinear features. Learning in this setting requires accounting for uncertainty in both the systematic evolution of the state and its stochastic fluctuations. We propose a reinforcement learning algorithm that jointly estimates these unknown components and incorporates model uncertainty into exploration. Our main result establishes sublinear regret with an explicit dependence on the diffusion strength, providing sharper guarantees in regimes with weaker stochastic disturbances without requiring prior knowledge of the diffusion. The analysis combines statistical estimation bounds with Itô’s lemma and the Fokker–Planck equation. Specifically, Itô calculus characterizes the stochastic errors in learning the dynamics, while the Fokker–Planck equation connects model estimation errors to discrepancies in state distributions and expected rewards. Together, these tools enable a diffusion-dependent analysis of cumulative regret, extending variance-aware learning guarantees to continuous-time control with unknown stochastic dynamics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.