Q-Guided Action-Transport Editing for Real-World Reinforcement Learning of Vision-Language-Action Models
Abstract
Flow-based vision-language-action (VLA) policies provide expressive priors for robot manipulation. Adapting them with online reinforcement learning remains difficult because each action is produced by a multi-step sampler. Backpropagating through the complete transport is costly, while an output-space correction cannot use the remaining generative dynamics to refine its edit. We introduce Q-Transport, an off-policy adaptation framework that inserts a stochastic local edit at an interior flow state. A conservative critic evaluates the resulting action, and its gradient reaches the edit policy through only the remaining frozen Euler steps. Q-Transport uses a fixed interior point produced by a stopped frozen prefix. Aggressive Q-Transport instead samples randomized interior points and directly corrupts completed actions, learning across different transport lengths. Both variants share an objective that combines action value, entropy, and optional prior-score regularization, and both select between paired base and edited candidates at inference. The edit location also unifies lightweight adaptation strategies: output-space residual methods act at the clean endpoint with an identity tail, initial-noise methods act at the noise endpoint before the complete sampler, and Q-Transport uses partial pretrained transport between these endpoints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.