acceptodds
Under review as a conference paper at ICLR 2027

Language Model Activations Inhabit Privileged Error-Correcting Basins

Abstract

Language models exhibit remarkable robustness, continuing to produce coherent text even when their activations are perturbed by interventions like linear steering. We hypothesize that this robustness is a result of passive dynamics, i.e., constraining mechanisms in the forward pass that funnel activations toward “good” regions that produce coherent outputs, similar to how biomechanical constraints allow living organisms to produce natural movement with minimal effort. To understand such constraining mechanisms in language models, we probe the geometry of their activation space by observing the action of model layers on low-dimensional curves. In doing so, we discover that model activations occur within a cluster of distinct attracting basins, which differentiate natural activations geometrically from distributionally similar synthetic activations. Applying this lens to language model steering, we observe feature-specific basins along semantic steering directions, and find that steering moves activations between these basins. To demonstrate the active role of this geometry in neural computation, we show that adaptively modulating steering strength to transport activations across basins improves inter-language steering, significantly increasing the probability of sampling tokens from the target language compared to fixed-strength steering. Our findings establish analysis of activation space geometry as a promising approach to interpreting language models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.