acceptodds
Under review as a conference paper at ICLR 2027

OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models

Abstract

Post training via GRPO has demonstrated remarkable effectiveness in improving the generation quality of flow-matching models. However, GRPO suffers from inherently low sample efficiency due to its on-policy training paradigm. To address this limitation, we present OP-GRPO, the first Off-Policy GRPO framework tailored for flow-matching models. First, we maintain a replay buffer that retains high-quality trajectories and reuses them by mixing a small fraction of off-policy samples into each rollout batch. Second, to mitigate the distribution shift introduced by off-policy samples, we propose a sequence-level importance sampling correction that keeps GRPO's clipping anchored to the current policy, reducing the clipped off-policy samples while ensuring stable policy updates. Third, we theoretically and empirically show that late denoising steps yield ill-conditioned off-policy ratios, and mitigate this by truncating trajectories at late steps. Across image and video generation benchmarks, OP-GRPO achieves comparable or superior performance to Flow-GRPO with only 34.2% of the training steps on average, yielding substantial gains in training efficiency while maintaining quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.