Edit Only Where It Fails: Boundary-Guided Editing via Latent Transformations for Vision-Language-Action Models
Abstract
Vision-language-action models provide strong priors for robot manipulation, but their performance can degrade under deployment-specific differences in object geometry, viewpoint, contact conditions, and robot embodiment. Existing adaptation methods often modify the policy globally, even when deployment failures are concentrated in a limited set of states. This creates a tension between adaptation and preservation: sufficiently strong updates may correct new failures but also disturb useful pretrained behaviors. We present BELT-VLA, a boundary-guided policy editing framework for online reinforcement learning of vision-language-action models. BELT-VLA identifies failure-prone regions through human interventions and applies state-conditioned orthogonal transformations to the action-flow representation only when an edit is required. The edit policy is optimized with an objective that balances exploration against intervention risk, while the orthogonal editing mechanism limits unnecessary policy deviation. We evaluate the method on four real-world manipulation tasks involving precise insertion, assembly, packing, and liquid transfer. The proposed approach increases the average success rate from 33.5% to 82.0%, while reducing human intervention and limiting policy changes outside failure-critical states. Comparisons with parameter-efficient fine-tuning and unconstrained adaptation show that boundary-guided orthogonal editing provides a favorable balance between online policy improvement and prior preservation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.