Improving the Robustness of VLAs to Spatial Object Perturbation via Attention-Head Amplification
Abstract
Vision-Language-Action (VLA) policies struggle under small perturbations, which limits robotic deployment. We focus on spatial object perturbation. When the target object is displaced from its training-time position, the policy fails to manipulate it and often reaches toward the training-time position. With activation patching we show that this failure reflects non-use of the object's position rather than its absence, and we propose a training-free method that mitigates it by amplifying the few attention heads that carry that position to the action tokens. On OpenVLA, OpenVLA-OFT, , , and GR00T N1.7 on LIBERO-Object, the new object location is present in a few object patches of the input visual tokens and is consumed in a limited band of layers. We then test every attention head one at a time. We run the policy on the scene with the moved object, but feed that head the output it would produce if the object were still at its training-time position, and measure how much the planned action stops following the object. We found that in every policy a small number of heads carry the displacement. Amplifying the input-dependent deviation of these heads raises paired closed-loop success, at 10 cm from 12.3% to 17.9% in OpenVLA-OFT and from 29.1% to 41.8% in , and at 15 cm from 42.1% to 50.5% in . The gain reproduces on the test states and transfers to LIBERO-Pro.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.