Spectral Blind Spots: Structure, Control, and Functional Effects in Language Model Heads
Abstract
Language-model output heads can develop highly concentrated singular spectra. We investigate what shapes this concentration, how it can be controlled, and when spectral intervention improves prediction. Under matched from-scratch training, spectral skew increases as the hidden-dimension-to-vocabulary ratio falls and hidden dimension grows. An empirical relation predicts held-out architectures within this regime, while pretrained-family comparisons expose its transfer limits. We introduce Spectral Skew Regularization (SSR), which penalizes a dominant-to-weaker singular-value ratio above an empirical threshold. Its pure-flow gradient lowers the dominant endpoint and raises the weaker one; with task gradients, we derive a sufficient condition for spectral descent. Experiments demonstrate spectral control across scales. Three-seed continuations at Pythia-410M and 1B improve on task-only training. SSR outperforms spectral-norm penalties on TinyStories at 410M, but SN performs better on both TinyStories and WikiText-2 at 1B. Prediction-preserving transformations and controlled perturbations show why lower skew alone does not imply better prediction. These results establish a conditional role for selective spectral intervention, with functional benefits that depend on the model and continuation setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.