acceptodds
Under review as a conference paper at ICLR 2027

Advantage-Gated Anchoring for Diffusion Reinforcement Learning

Abstract

Advantage-weighted regression enables diffusion models to learn from reward feedback, providing a practical approach to diffusion reinforcement learning. For samples with positive advantages, regression encourages fitting their associated velocity targets. However, negative advantages push predictions away from these targets, without reallocating probability mass to alternative outcomes as negative feedback does in policies that explicitly predict action probabilities. This motivates studying policy anchoring as an additional constraint on velocity updates, acting as a reference for learning from negative feedback. In this work, we analyze how anchor policies shape reward learning and training stability in advantage-weighted regression and reveal that anchoring helps stabilize negative-advantage updates, while applying it to positive-advantage samples can hinder reward learning. Based on these insights, we propose Advantage-Gated Anchoring (AGA), a simple yet effective method that anchors only negative-advantage samples. Experiments demonstrate faster reward improvement, a better balance between reward and image quality, and greater robustness to the choice of anchor policy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.