acceptodds
Under review as a conference paper at ICLR 2027

One-Step Steering is Not Sequential Control

Abstract

Sparse autoencoder (SAE) features give steering methods activation directions, but a direction specifies only one step: once the first intervention moves the hidden state, the second must be evaluated at a state where it was never specified. We call the rule that resolves this the execution law and study it for encoder–decoder controllers. We derive an exact identity showing that re-encoding after each step adds an interaction defect driven by the encoder–decoder round trip, and prove that good reconstruction does not bound this defect. Across eight public GPT-2 SAE widths, wider dictionaries reconstruct better while their round-trip error grows monotonically. On 493 held-out feature pairs in GPT-2 and Gemma-2-2B, the identity predicts which application order produces less interaction (Spearman 0.98–0.99), although accuracy falls when a controlled feature changes support. Choosing the predicted order cuts the interaction residual by 24–38%, yet direct error to the joint target changes by about 1% or less, because interaction accounts for only 4–7% of endpoint error and can cancel singleton errors. Reconstruction, interaction fidelity, and target delivery are therefore separate evaluation axes, and sequential steering must specify its execution law and measure outcomes directly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.