Geometry Aware Diffusion Policies for Online Reinforcement Learning
Abstract
Diffusion models have emerged as a powerful class of expressive policies in Reinforcement Learning (RL). In this paper, we propose an improved diffusion policy method for online RL, where the diffusion process is driven entirely by a learned Q-function and adapts to the geometry of the action-value landscape. Unlike conventional diffusion policies that rely on separately trained score networks, our formulation approximates the score function by the gradient of the Q-function, motivated by the Boltzmann exploration. Furthermore, we introduce a state-dependent preconditioner, motivated by the geometry of Q function, using the covariance of Q-gradients across parallel sampling chains without additional network evaluations, which yields structured-anisotropic exploration grounded in information geometry. The algorithm is evaluated on the Gymnasium benchmarks and attains the highest aggregate normalized return among diffusion reinforcement learning baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.