Time Scales in Diffusion Policy Learning: Action-Gradient Attenuation and Poisson-Based Policy Improvement
Abstract
Continuous-time diffusion reinforcement learning involves two distinct time scales: the evolution of the environment and the dynamics used to sample actions. We show that a mismatch between these time scales changes the critic signal used to update the actor. Under stationary action resampling, the critic action-gradient signal vanishes linearly with the physical time step \(h\), leading to increasingly weak actor updates as \(h\) tends to 0. We derive a finite-horizon corrector that connects this finite-step effect to the stationary-resampling and fast-action limits, and characterize the continuous-time policy-improvement direction through a performance-difference identity based on the Poisson equation. Normalizing the finite-step actor signal by \(h\) restores its magnitude, but recovers the limiting policy-improvement direction only when the relevant action modes share a common relaxation rate. An exactly solvable linear–quadratic (LQ) model further shows that a fixed scalar normalization can improve the first policy update yet cause the policy updates to diverge when applied repeatedly. Experiments with both exact and learned critics validate the predicted linear attenuation and show that normalization by \(h\) improves learning at small physical time steps, while performance remains sensitive to the choice of scalar coefficient.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.