SPECTRAL SHAPE RESIDUAL: A LIGHTWEIGHT SPECTRAL DIAGNOSTIC FOR ATTENTION WEIGHT MATRICES
Abstract
We introduce the , a lightweight metric computed directly from the query and key weight matrices of attention layers - no forward pass, no training data needed. SSR measures how similar the singular value distributions of and are: when they match perfectly, SSR is zero. We study SSR across four settings organized as a analytical framework (toy vs. production LLM; static snapshot vs. training dynamics) - a design that systematically covers the conditions under which spectral weight-space signals can arise. In toy transformers trained on modular arithmetic, SSR falls sharply during training, dropping below 50% of its peak value 200-1,900 steps the model generalizes - making SSR an early warning signal - and converges to near-machine precision once grokking is complete. In production LLMs (LLaMA-3-8B, Gemma-4-31B-IT), SSR decreases systematically with layer depth, and RL-distilled models (DeepSeek-R1-Distill-Qwen-14B) show lower SSR than their base model (Qwen2.5-14B-Instruct) in 45 out of 48 layers. In OLMoE-1B-7B pretraining, 18.8% of attention heads show persistent monotonic decline in SSR before grokking completes, concentrated in deep layers (L10–L13); the median SSR of these heads falls 40% below its peak 582B tokens before grokking completes. SSR requires only one SVD per attention head and no evaluation data, making it deployable as a zero-overhead structural health monitor directly inside training pipelines. For both pre-training (does the model grok?) and post-training (does RL alignment shift spectral structure?), SSR provides a continuous signal without running a single forward pass.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.