acceptodds
Under review as a conference paper at ICLR 2027

Transition-GRPO: Improving Flow Model Alignment with Direct Transition Targets

Abstract

Reinforcement learning has emerged as an effective approach for aligning diffusion and flow-based generative models with downstream preferences. Backward-process methods such as Flow-GRPO optimize stochastic reverse-SDE transitions, but often exhibit slower reward improvement than forward-process alternatives, raising the question of how these transitions can be optimized more effectively. We revisit the Flow-GRPO objective and show that its local update reveals a simple policy-improvement direction in transition-mean space. Building on this observation, we propose Transition-GRPO (T-GRPO), which uses a KL-proximal formulation to derive explicit, advantage-adjusted transition-mean targets and replaces likelihood-ratio optimization with direct regression toward these targets. We further extend this formulation to jointly perform CFG distillation and reward optimization during a short initialization stage, using CFG-generated trajectories to construct advantage-adjusted transition targets for a non-CFG model. Experiments demonstrate that T-GRPO consistently accelerates convergence and achieves higher rewards under a fixed training budget than Flow-GRPO. It also surpasses the forward-process DiffusionNFT baseline on preference-based rewards, while DiffusionNFT remains stronger on structured tasks. Moreover, T-GRPO yields additional gains when integrated with existing acceleration methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.