acceptodds
Under review as a conference paper at ICLR 2027

Attention Invariants: Weight-Space Interpretability Without a Forward Pass

Abstract

Attention heads in transformers can be characterized through basis-invariant quantities from weights alone, without any data or forward pass. We derive two invariants of the QK sector: , which measures a head's positional capacity, and , the definiteness of the symmetric part of its QK form. We prove a necessity condition on positional computation. A head with cannot depend on relative position, so positional function requires , a structural constraint we find is obeyed without exception across every model. On five models from four architecture families (Qwen2.5-1.5B/7B, Mistral-7B, Llama-3.1-8B, Gemma-2-2B), and separate three canonical head types (induction, duplicate-token, previous-token) in distinct regions of a two-dimensional plane. We show that a logistic regression trained on separates every head type pair at ROC-AUC , and unsupervised clustering recovers the types at – purity. Transferring to a different architecture, we obtain a median of balanced accuracy ( of pairs ). As an application, we apply YaRN only to the top heads by , arriving at better results than a standard uniform application.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.