SPARSE TRANSCODERS FOR MECHANISTIC ANALYSIS OF VLA POLICIES
Abstract
We adapt sparse time-conditioned transcoders (STCs) to the action expert of the flow-matching vision-language-action policy π0.5, giving a sparse intervention in- terface for systematically studying continuous action prediction. We train one STC for each of the 18 action-expert MLP blocks; the resulting sparse codes activate only 5.02% of the dictionary on average and support feature discovery, circuit tracing, open-loop ablation, and steering. In a discovery-selected black- bowl / bottom-drawer case study, a rare feature firing on only 0.166% of probed observations yields a traced 16-node upstream circuit whose ablation produces about a 2× larger action-chunk change than ablating the target alone, and about a 2× larger effect than matched same-size random circuits in the evaluated cir- cuit setting. Target ablation is zero on negative controls where the feature is in- active, while behavior-grounded steering increases the Cartesian-speed score on 100% of evaluated examples for target-feature steering and 85% for circuit steer- ing; an agreement-filtered contrastive direction changes the gripper coordinate by about 1.6 at α = 2, compared with about 0.05 for a matched random direc- tion. Separately, we run a visual-counterfactual audit using non-time-conditioned TopK transcoders on the PaliGemma language-prefix MLPs. There, 83.72% of a capped set of visually selective features reproduce their selectivity on held-out frames, but small sparse patches have limited and direction-dependent action ef- fects, and attribution-based action restoration does not improve closed-loop suc- cess over occlusion-only or random-restoration controls. Together, the two stud- ies show that sparse transcoders can support feature and circuit interventions in VLA policies while separating feature selectivity, open-loop action effects, steer- ing, closed-loop behavior, and evidence that a circuit carries its attributed visual content
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.