RAM-A: Bounded Advantage Regression for Flow Policy Optimization
Abstract
Flow-matching policies provide an expressive class of action distributions for robotic control, but reinforcement learning with such policies remains challenging. Existing approaches such as Flow Policy Optimization (FPO) construct policy updates through exponentiated differences of flow-matching losses and rely on proximal clipping to stabilize optimization. We introduce RAM-A, an advantage-based extension of Reinforce Adjoint Matching (RAM) that casts reinforcement learning as regression toward advantage-shifted flow-velocity targets. We show that RAM-A and FPO share the same first-order policy-improvement direction at the behavior policy, but differ under repeated optimization of a rollout batch, where unbounded advantage regression diverges for negative advantages. RAM-A instead bounds the norm of the regression target, which makes each update self-limiting and robust to corrupted advantage estimates. We evaluate RAM-A across dense-reward continuous control and sparse-reward robotic manipulation, together with targeted experiments on multimodality and multitask retention. While matching existing methods in task performance, RAM-A retains both modes of a multimodal pretrained policy and better preserves skills not targeted by the fine-tuning reward. Pretrained generative policies can thus be improved one skill at a time without erasing their diversity or capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.