DS-GRPO: Noise-Space Adaptation of Diffusion Policies with Action-Space Verification
Abstract
Diffusion steering adapts a pretrained generative policy without altering its weights: a small steering policy chooses the initial noise of the frozen sampler, re-weighting the base policy's behaviors within their support. Yet the signals one could adapt to, e.g., a demonstration, a success, a distance to a target, are all only verifiable in action space, while steering optimizes in noise space. Existing methods resolve this mismatch with different trade-offs. Critic-based noise-space RL verifies in action space but pays the cost of exploration and value estimation through environment transitions. Inversion-based steering instead maps demonstrated actions back to noise space, regressing on noise labels rather than minimizing action-space distance. We propose **DS-GRPO**, combining the benefits of both supervision from demonstrations and direct verification in action space. DS-GRPO samples a group of noises from the steering policy, decodes them through the frozen sampler, and uses the demonstration as an action-space verifier to score the resulting actions. These scores guide steering-policy updates through group-relative policy gradients. Across simulation and real-world experiments including offline behavior cloning, goal-conditioned behavior cloning, and online DAgger, DS-GRPO matches or outperforms pre-existing steering methods on most tasks, providing a concrete step towards efficient yet effective policy adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.