acceptodds
Under review as a conference paper at ICLR 2027

Bridging the Deployment Gap in Diffusion Reinforcement Learning via State-Conditioned Latent Control

Abstract

Recent diffusion and flow policies provide expressive action distributions and achieve strong performance in reinforcement learning, yet their deployment on physical robots remains underexplored. We identify a key deployment gap: although diffusion policies outperform Gaussian policy baselines in simulation, independently resampling latent inputs at every control step introduces temporal inconsistency, leading to high-frequency action variation, increased body sway, and trajectory dispersion. Fixing the latent input reduces this variation but cannot adapt latent behaviors to changing robot states and task commands. To address this issue, we introduce SteerGenPO, a state-conditioned latent control framework that freezes the reinforcement-learning-trained generator and learns an actor to select latent inputs from the robot state and task command. An angular parameterization constrains the latent input to a sphere of radius , matching the typical radial scale of the Gaussian prior used during generator training. During deployment, the actor deterministically selects latent inputs using its mean angular coordinates. We evaluate SteerGenPO on robot control benchmarks, Unitree G1 locomotion and path tracking, and physical-robot experiments. Compared with random latent sampling, SteerGenPO reduces high-frequency action variation and trajectory dispersion while improving command tracking, enabling more reliable deployment of diffusion reinforcement learning on real robots.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.