acceptodds
Under review as a conference paper at ICLR 2027

Flow Policy Optimization with Steerable Editing

Abstract

Flow policies offer an expressive and scalable foundation for continuous control, but online reinforcement learning becomes subtle when generated actions are refined before execution. An action with low current value may still be a valuable sample if a small edit can move it into a high-value region, so the generator should account for what its actions can become after correction. We introduce Steerable Flow Editing (SFEdit), a proximal framework that derives distribution learning and action refinement from a single policy-improvement objective. A Wasserstein-smoothed KL regularizer yields an exact tilt–transport factorization: the flow favors actions by the value attainable after correction, net of the movement cost, and a proximal map performs the correction. The scaled proximal displacement provides the actor learning signal, and the same map determines the actions used for interaction and critic backups. We realize the resulting tilt through adjoint matching and introduce iQAM, an integrating-factor discretization that reduces the degradation of few-step learning without additional vector–Jacobian products. On 50 OGBench tasks across ten domains, SFEdit achieves 87.6% final aggregate success, compared with 80.7% for the strongest baseline under the shared evaluation protocol. On two fine-grained real-robot manipulation tasks, it nearly triples the success of a frozen generalist within 45 minutes of online training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.