acceptodds
Under review as a conference paper at ICLR 2027

Not All Concepts Are Causal: Sparse Autoencoders Reveal a Motor Layer in VLAs

Abstract

Vision-language-action (VLA) policies hide which internal features merely describe a task and which produce the action. We introduce SAE-VLA, which turns a frozen VLA's activations into a concept lookup table for reading out and steering its actions. A sparse autoencoder provides the rows, a vision-language model names them, and a blinded verifier checks every name on held-out episodes. A signed intervention records whether a row moves the action its name implies, at a dose certified by a closed-loop gate. Across five policies on two backbones, where the table is built decides whether it works. In , language features receive plausible motor names that verify half as often and barely act. The action expert instead holds a motor layer of grasp, release, hand-height, and wrist-tilt features. On the breakfast policy, verified names identify the moving arm on of held-out single-arm frames and contain the current subtask on , by lookup alone. A name indicates where to steer but not whether, so the table's causal column defines what is steerable. On the LIBERO policy, a gated verified row makes the arm descend earlier with success unchanged. Gated handles replicate with a second dictionary and on the OpenVLA-OFT LIBERO policy. Steering by language lets the frozen breakfast policy recover bread placements beyond its demonstrations, raising task success from to without new data or retraining. Not every named concept is causal; the table records which ones are.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.