RLFR: Extending Reinforcement Learning for LLMs with Flow Environment
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising framework for strengthening reasoning abilities of Large Language Models (LLMs). While policy optimized with binary verification prone to overlook potential valuable exploration in reasoning trajectory. Given the heavy annotation cost of Process Reward Models (PRMs), recent works investigate auxiliary signals of reward shaping for trajectory tokens, including entropy, likelihood, and teacher divergence derived from the logit space. In this work, we offer a novel perspective on shaping RLVR with flow rewards derived from latent space, and propose RLFR, where the flow fields of model latents are constructed from either off-policy high-quality data and on-policy rejection sampling data, and the velocity deviations of policy latents within it are quantified to serve as a reward signal. RLFR first demonstrates that a well-established flow field can be a sound environment for reward signal collection, highlighting the expressive yet context dependent latent space is much underexplored. Moreover, RLFR is able to compress any off-policy expert data as reference for constituting reward signals, and we show that RLFR achieves efficiency on par with RLVR for additional trajectory rewarding. Experiments on both language and multimodal reasoning benchmarks demonstrate the reliability of flow rewards, and suggesting a promising paradigm for credit assignment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.