acceptodds
Under review as a conference paper at ICLR 2027

Beyond Target Deltas: Validating Activation Steering as a Model Intervention

Abstract

Activation steering is often evaluated through a change in target score. We examine how the fitted vector, the behavioral readout, and the generated text affect the interpretation of that change, using 105,000 generation records from seven model variants and a separate vector-extraction audit. With source examples and estimator fixed, raw-text and chat-template extraction yield mean cross-layer cosine similarities of 0.325 for Llama-3.1 and 0.309 for Qwen2.5. On fixed sycophancy responses, factual correction falls by 19.6 pp even as generic compliance and false-premise endorsement rise. In a Mistral jailbreak cell from the 50-pair α×k grid, lower harmful compliance accompanies a 96 pp increase in degraded outputs. Four of the 21 primary configurations meet the joint control–degradation cutoffs by point estimate, but none has confidence bounds satisfying both cutoffs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.