Amplification Without Repair: Causal Limits of Activation Steering in Abstract Reasoning
Abstract
Activation steering modifies hidden representations to control language-model behavior. Current approaches select interpretable features or learn directions from demonstrations, seeking improvements in reasoning. How these interventions translate into correction on new examples remains unclear. We investigate this transfer through a mechanistic study of text-serialized ARC-AGI-1, ARC-AGI-2, and symbolic Bongard-LOGO. Feature audits combine input controls, activation localization, and random-controlled interventions; support-adaptation experiments examine gradient alignment, decision boundaries, and support aggregation. Steering improves output validity and produces concept-specific held-out margin gains, while decision repair remains sparse. Fit/check gradient alignment differs between initially correct and incorrect concepts across development and fresh-combination confirmation. Aggregating support gradients amplifies both groups' responses, strengthening correct and incorrect preferences. These findings connect correction to directional transfer and boundary crossing, motivating interventions that target held-out errors and evaluations that track repairs and harms alongside margin gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.