acceptodds
Under review as a conference paper at ICLR 2027

RAM-A: Bounded Advantage Regression for Flow Policy Optimization

Abstract

Flow-matching policies provide an expressive class of action distributions for robotic control, but reinforcement learning with such policies remains challenging. Existing approaches such as Flow Policy Optimization (FPO) construct policy updates through exponentiated differences of flow-matching losses and rely on proximal clipping to stabilize optimization. We introduce RAM-A, an advantage-based extension of Reinforce Adjoint Matching (RAM) that casts reinforcement learning as regression toward advantage-shifted flow-velocity targets. We show that RAM-A and FPO share the same first-order policy-improvement direction at the behavior policy, but differ under repeated optimization of a rollout batch, where unbounded advantage regression diverges for negative advantages. RAM-A instead bounds the norm of the regression target, which makes each update self-limiting and robust to corrupted advantage estimates. We evaluate RAM-A across dense-reward continuous control and sparse-reward robotic manipulation, together with targeted experiments on multimodality and multitask retention. While matching existing methods in task performance, RAM-A retains both modes of a multimodal pretrained policy and better preserves skills not targeted by the fine-tuning reward. Pretrained generative policies can thus be improved one skill at a time without erasing their diversity or capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.