Few Neurons, Many Roles: Sparse MLP Neurons That Control Attention Sinks
Abstract
Large language models (LLMs) send a disproportionately large share of attention to the first token, a pattern known as the attention sink. Although this pattern is widespread and functionally important, it is debated which upstream units build it and whether those units serve only the sink. We identify sparse populations of MLP neurons that causally control the sink across two model families. Suppressing these neurons weakens the sink by 12 to 51%, while amplifying their natural activity strengthens it by 1 to 6%. Exact softmax decomposition and live restoration show that these neurons act through later queries and keys, with a pathway mix that differs across models. Evaluating the models with the neurons suppressed reveals that these neurons significantly change downstream behavior differently across model families and benchmarks. Restoring the exact sink while the neurons remain suppressed recovers little or none of the behavioral effect, indicating that the sink explains little of this effect We further find that preserving their coordinates in high precision while quantizing the model to W8A8 recovers 65% of the loss, despite these neurons having little overlap with magnitude-selected neurons. A few neurons help build the sink, but their role reaches well beyond it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.