acceptodds
Under review as a conference paper at ICLR 2027

Amplification Without Repair: Causal Limits of Activation Steering in Abstract Reasoning

Abstract

Activation steering modifies hidden representations to control language-model behavior. Current approaches select interpretable features or learn directions from demonstrations, seeking improvements in reasoning. How these interventions translate into correction on new examples remains unclear. We investigate this transfer through a mechanistic study of text-serialized ARC-AGI-1, ARC-AGI-2, and symbolic Bongard-LOGO. Feature audits combine input controls, activation localization, and random-controlled interventions; support-adaptation experiments examine gradient alignment, decision boundaries, and support aggregation. Steering improves output validity and produces concept-specific held-out margin gains, while decision repair remains sparse. Fit/check gradient alignment differs between initially correct and incorrect concepts across development and fresh-combination confirmation. Aggregating support gradients amplifies both groups' responses, strengthening correct and incorrect preferences. These findings connect correction to directional transfer and boundary crossing, motivating interventions that target held-out errors and evaluations that track repairs and harms alongside margin gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.