Reshaping Action Flow: Behavior Correction for Flow-Matching Vision-Language-Action Policies
Abstract
Vision-language-action (VLA) models have shown remarkable capability across diverse manipulation tasks. After deployment, however, the pretrained policy may face situations unseen in training, for example an unexpected obstacle on its path it must manage to avoid on its own. Current methods either intervene at inference time, so the policy itself never learns the new behavior, or fine-tune the policy on new demonstrations, a route known to disturb the policy's original ability. In this paper, we propose a behavior correction framework for flow-matching VLA policies. Given a few correction demonstrations, it learns the correction through a guided training target and maintains the model's original ability through a gradient projection method. We further exploit the locality of the correction, giving more protection to the frames with larger correction and less to the rest, so the update interferes less with the original ability. Extensive experiments on robot manipulation verify the effectiveness of our method in correcting a deployed policy while preserving its original ability, with locality further strengthening the protection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.