Observing and Steering Affordance in VLA Hidden States for Spatial Robustness
Abstract
Vision-language-action (VLA) policies achieve strong performance on manipulation benchmarks, but their success rate degrades sharply once object positions drift outside the training distribution, limiting real-world deployment. We ask whether a VLA's internal representations already encode a high-level semantic feature that reflects task progress, and whether this feature can be observed and controlled at inference time, without retraining the policy. We show that a linear probe trained on language-backbone hidden states reliably decodes an affordance-based cost derived from 3D value maps, and that this cost can be driven toward a target value at test time through targeted activation perturbations along the probe's gradient direction. A layer-wise causal analysis identifies intervention layers where this perturbation both lowers the predicted cost and changes the resulting action, which we use to select where steering is applied for each task. We evaluate this observe-and-control approach on two recent open-source VLA backbones, OpenVLA and , under a simulated closed-loop out-of-distribution benchmark built on LIBERO-Pro, in which the target object is displaced by an increasing radius from its training position. Without modifying any policy weights, steering raises out-of-distribution task success from 16.6% to 29.5% on OpenVLA and from 34.5% to 41.1% on , an average absolute improvement of nearly 10 percentage points across both backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.