Not All Errors Are Equally Steerable: Predicting Error Types to Guide Activation Steering
Abstract
Activation steering aims to correct language-model errors by modifying internal activations at inference time, but different errors may not be equally amenable to the same intervention. We ask whether layer-wise token trajectories can reveal which errors are most likely to benefit from steering. Tracing the correct-answer token across layers reveals two recurring patterns. In a no-entry error, the correct token never becomes a plausible output candidate; in a competitive-decay (CD) error, it becomes competitive at an intermediate layer but loses its advantage before the final output. Under the same fixed steering method, CD errors are recovered approximately four times as often as no-entry errors. Motivated by this gap, we introduce Information Dynamics, a framework for characterizing these trajectories and predicting error type. Although the taxonomy is defined using the reference answer, error type can be predicted from the model’s own early candidate dynamics without access to the reference answer at inference time. Across six open-weight models and eight tasks, these dynamics consistently distinguish CD from no-entry errors. We use these predictions to allocate a fixed steering budget toward more recoverable errors. Across seven tasks on Qwen3-8B and Llama-3.1-8B, steering only 20% of errors selected by predicted CD probability corrects more than twice as many errors as random selection under the same intervention budget. These results show that layer-wise token dynamics provide a practical signal of error steerability and can substantially improve activation steering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.