Geometry-Adapted Stepping for Stochastic Control via Walk-On-Spheres
Abstract
Reinforcement learning (RL) algorithms with a continuous action space rely on discrete (Euler-Maruyama) time-stepping to simulate continuous-time stochastic dynamics, introducing temporal discretization error and requiring ad hoc collision handling to respect geometric boundaries. These issues may be improved with smaller integration steps, but this increases overall simulation cost, creating a tradeoff between performance and speed. We introduce Walk-on-Spheres (WoS) Monte Carlo rollouts that perform geometry-aware stepping to traverse the state space more efficiently, are collision-free by construction, and do not require time discretization. In the setting of quadratic control costs, relevant to KL-regularized reinforcement learning, we utilize the Doob -transform to learn an optimal control policy for fixed and goal-conditioned navigation problems with only terminal rewards. We evaluate our method on 2D point-maze and 3D HM3D indoor navigation scenes, demonstrating significant speedups in rollout time while exceeding baseline performance under equal compute budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.