acceptodds
Under review as a conference paper at ICLR 2027

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO

Abstract

Group Relative Policy Optimization (GRPO) has demonstrated remarkable success in aligning text-to-image (T2I) flow models with human preferences. However, we have identified that the learning loop of current flow-based GRPO is fundamentally decoupled from the learner's current capability, suffering from critical blind spots at both prompt selection and advantage estimation: (i) Existing methods sample prompts randomly, overlooking the substantial impact of data selection on reinforcement learning (RL) efficacy–a factor proven crucial in GRPO for large language models; (ii) They evaluate sample quality solely relying on intra-group statistics, lacking a global perspective to reliably assess policy progress. To address these issues, we propose Adaptive GRPO (AdaGRPO), a novel capability-aware RL algorithm tailored for flow models. Specifically, AdaGRPO consists of two principal components: (i) Online Curriculum Filtering Strategy dynamically tracks the model’s proficiency and adaptively selects prompts that best match its current learning boundary; (ii) Cross-Level Advantage Fusion synergistically integrates fine-grained intra-group advantages with global advantages, providing a comprehensive and globally calibrated sample-level evaluation. As a lightweight, plug-and-play module, AdaGRPO can be seamlessly integrated with existing frameworks such as Flow-GRPO, DanceGRPO, and Flow-CPS. Extensive experiments demonstrate that AdaGRPO consistently drives performance gains while significantly stabilizes GRPO training for flow models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.