Paired Residual Adaptation for Vision-Language-Action Policies under Transient Visual Mislocalization
Abstract
A brief error in a target's apparent position can redirect a robot's actions and cause a placement to fail, even after the visual error disappears. We study how to reduce these failures through action-level adaptation of a frozen vision-language-action (VLA) policy. Our approach learns a lightweight residual action policy from paired predictions of the base model. Clean and visually perturbed observations of the same physical state are evaluated with matched sampling noise, and their action difference provides the correction target. Balanced zero-residual supervision discourages unnecessary changes to valid behavior. At deployment, the residual policy uses current images, robot state, and the proposed action chunk to adjust translational actions, without a clean reference or additional base-model queries. We evaluate the method with a frozen π0.5 policy on LIBERO-Spatial, using brief rendering-only target displacements during placement. A fixed residual checkpoint improves task success on held-out initial states. Small development comparisons under matched training budgets also favor the paired residual design over direct action prediction, mismatched references, and unbalanced supervision. These results support targeted action adaptation for transient visual errors, although some clean-condition performance is lost and broader generalization remains to be established.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.