acceptodds
Under review as a conference paper at ICLR 2027

Beyond the KL Brake: Reward-Level Adaptive Anchoring for Diffusion Policy Optimization

Abstract

Diffusion models are difficult to directly align with specific preferences. While Reinforcement Learning (RL) improves alignment, it suffers from sparse rewards. A common method is to assign the only final reward to all timesteps, but this weakens credit assignment. When rewards are identical across steps, action-level differences are obscured, making it hard to correct erroneous denoising behaviors and increasing the risk of distributional shift. Kullback–Leibler (KL) divergence is widely used to prevent distributional collapse, yet the KL-as-loss paradigm has limitations. First, it only indirectly regularizes parameters and cannot directly suppress harmful actions. Local drift caused by such actions can accumulate through diffusion’s iterative process, a problem observed in UNet-based RL and amplified in Flow RL with SDE noise. Second, under KL-as-loss, the model must continually overcome the optimization resistance introduced by a static regularization term while pursuing higher rewards. As a result, more time is required to reach the same level of performance as training without KL. Third, a fixed KL weight creates a trade-off between preserving the pretrained distribution (exploitation) and efficient alignment (exploration): small weights cause rapid distribution drift, while large weights over-anchor the model to the pretrained distribution, weakening preference optimization. To address the above issues, we propose \ourmethod, which converts KL-as-loss into fine-grained step rewards, making distributional deviation an explicit per-step feedback signal. This suppresses drift accumulation and mitigates reward sparsity. By jointly modeling exploration intensity and reward trends, \ourmethod adaptively adjusts anchoring strength, preserving pretrained advantages while improving alignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.