acceptodds
Under review as a conference paper at ICLR 2027

Where Massive Activations Come From: Causal Evidence from MLP Geometry in Large Language Models

Abstract

Large language models (LLMs) commonly exhibit massive activations (MAs) across Transformer architectures. MAs are sparse hidden-state coordinates whose magnitudes greatly exceed typical activations and concentrate at specific tokens and dimensions. These extreme values complicate quantization, compression, and efficient inference, yet their generation mechanism remains unclear. Existing studies mainly associate MAs with attention sinks or mitigate them as numerical outliers. However, these approaches do not explain how MAs are generated inside the network. In particular, the causal connection between token-level triggers, layer-wise sources, and weight-space geometry remains unresolved. We address this gap through a systematic study of 26 language models combining component ablations, singular value decomposition, and causal interventions. Our results show that MAs primarily arise from MLP down-projections, where function-token representations align with load-bearing right-singular directions that amplify and write extreme activations, while attention mainly regulates this pathway. These findings establish MAs as structured geometric events and provide a causal account of their emergence across architectures.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.