acceptodds
Under review as a conference paper at ICLR 2027

MRFT: MeanFlow Reinforcement Fine-Tuning via Tractable Stochastic Transitions

Abstract

Flow matching has shown strong performance in continuous robotic control, but its multi-step sampling procedures incur substantial rollout costs in online reinforcement learning (RL). MeanFlow presents a compelling solution to this sampling overhead by enabling efficient few-step generation through explicitly modeling the average velocity field. However, directly applying MeanFlow to online RL is hindered by a critical limitation: its deterministic sampling process lacks the stochastic exploration and tractable transition likelihoods required for RL fine-tuning. Existing techniques depend on auxiliary learnable networks for post-hoc stochasticity injection, inherently increasing computational overhead and parameter complexity. Rather than relying on heuristic designs, we formally derive MeanFlow-SDE from the underlying flow dynamics. By approximating the finite-interval integration of the marginal-preserving flow SDE, MeanFlow-SDE yields tractable Gaussian transitions that inherently support stochastic exploration, eliminating the need for additional components. Building upon MeanFlow-SDE, we introduce MeanFlow Reinforcement Fine-Tuning (MRFT), an efficient online fine-tuning framework tailored for MeanFlow policies. By embedding our tractable transitions into RL objectives, MRFT seamlessly integrates stochastic generative sampling with robust policy optimization. Extensive evaluations on lightweight control policies, -style architectures, and real-world robotic systems demonstrate consistent improvements in sample efficiency and task success rates while maintaining few-step inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.