Distributed Sparse Interventions in Language Models
Abstract
Language models can perform a wide range of tasks at various levels of abstraction. They flexibly infer tasks from context, execute multiple tasks simultaneously, and select among competing goals. Understanding and controlling how models represent these tasks is central to aligning them with human objectives. Researchers can change model behaviour in many ways, but model steering, which manipulates intermediate task representations, links representation to behaviour most directly. Current steering methods typically intervene densely and assume linear effects, e.g., shifting every activation along a global direction in the residual stream. Because such interventions change thousands of neurons at once, they leave researchers to separate relevant from irrelevant neurons. We introduce Distributed Sparse Interventions (DSI), which steers models through a sparse set of individual neurons distributed across layers and attention heads. Finding such a set requires accounting for the selected neurons' interactions, even across layers; DSI achieves that by iteratively refining which neurons to intervene on and how strongly. Across 12 tasks and three instruction-tuned models, DSI activates task behaviour by identifying and intervening on fewer than 0.1% of neurons, outperforming dense steering vectors and often matching 10-shot in-context learning. Because DSI identifies explicit sets of neurons, it connects neuron-level computation to changes in latent representations and, ultimately, to desired outputs, allowing us to study how distributed computational components interact to produce steered behaviour. Our results demonstrate that steering can be viewed not only as a linear shift in representation space, but also as a targeted intervention on model computation, realised by adjusting a small set of interacting neurons whose coordinated activity spans multiple layers and contributes to behaviour.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.