Steering Policies with Emergent Registers
Abstract
Vision-language-action (VLA) policies acquire rich manipulation skills, yet deployment may require using these skills in ways not seen during training, such as manipulating new objects, correcting an ongoing reach, or avoiding a designated region. Test-time steering can provide guidance to the policy for these issues without retraining. Prior steering methods largely modify the action-generation process, such as its latent noise or flow dynamics. We instead ask whether steering can act through visual representations within the VLM backbone. Across multiple VLAs, we find that action tokens concentrate attention on visually uninformative image patches, resembling registers in Vision Transformers. We call these action registers. Surprisingly, masking attention to action registers changes predicted trajectories far more than masking attention to task-relevant image tokens. We therefore introduce RegiSteer, which applies test-time residuals to action registers to steer frozen VLAs toward desired Cartesian directions while leaving detailed manipulation to the policy. RegiSteer improves manipulation across generalist and specialist policies without weight updates, remains robust to noisy spatial guidance, and supports spatial safety constraints.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.