acceptodds
Under review as a conference paper at ICLR 2027

Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition

Abstract

Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by learning long-range transport maps between noise and data. However, their deterministic transitions do not directly provide the stochastic trajectories and tractable likelihood ratios required by reinforcement learning (RL) post-training. Existing SDE-based stochasticization techniques target velocity-based samplers and do not directly extend to long-range flow-map transitions. We propose Flow-Map GRPO, an online RL post-training framework for deterministic few-step flow-map generators. Its key component, Anchored Stochastic Flow Map Composition (ASFMC), combines deterministic transport with anchor-based conditional resampling. We establish the conditions under which this construction preserves the marginal probability path and develop tractable local- and endpoint-anchor policies for two-time and single-time flow maps. These policies enable a unified GRPO training procedure. Experiments on FLUX-based MeanFlow and sCM generators demonstrate substantial improvements in OCR, PickScore, and GenEval at different numbers of inference steps, including joint OCR–PickScore gains with mixed rewards. Controlled ablations show that the stochastic transition design is essential for translating training rewards into generation quality. Flow-Map GRPO enables effective RL alignment of pretrained deterministic flow-map generators while retaining their original parameterization, without retraining them as native stochastic models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.