acceptodds
Under review as a conference paper at ICLR 2027

Sparse Feature Analysis and Control of Vision-Language-Action Models under Visual Perturbations

Abstract

Vision-Language-Action (VLA) models have demonstrated strong manipulation performance, yet often degrade under visual perturbations such as camera, background-texture, lighting, and sensor-noise shifts. We ask which internal features make a VLA visually sensitive, what these features represent, and whether intervening on them produces the predicted change in behavior. To obtain localized analysis units, we decompose raw VLA activations using sparse autoencoders (SAEs), and introduce a shared-state paired perturbation protocol that isolates feature responses to visual shifts while controlling for physical state divergence. We then identify compensator features, which serve as action-coupled steering targets, and detector features, which predict perturbation-induced failures. Across autoregressive and diffusion-based VLA architectures, sparse features provide more localized perturbation-relevant units than raw residual dimensions, and the two roles form disjoint populations: detectors in the vision-language stream or early layers and compensators in the action expert or late layers. By relating features to the policy’s own action modes, we find that compensators encode task-invariant motion primitives and that a perturbation weakens the policy’s commitment to the primitive it is executing rather than switching it to another. Finally, we verify the interpretation by intervention: steering a single compensator feature shifts perturbed actions back toward the unperturbed action and modestly improves rollout success, and gating it with a detector makes it perturbation-specific. Together, these results identify, interpret, and verify the features behind visual sensitivity in VLAs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.