acceptodds
Under review as a conference paper at ICLR 2027

-Align: Assessing Language Model Representations Through Prediction-Head Alignment

Abstract

Representations with identical spectra can lead to very different logits, depending on their relationship with the model's prediction head. Existing spectral metrics describe how representational variance is distributed, but overlook how the model transforms the variations into predictions. We introduce -Align, a label-free metric that quantifies the alignment between representational variance and the directional sensitivity of the language modeling head. By combining the complete representation spectrum with the head's response along each principal direction, -Align captures a property that representation-only metrics ignore. Experimental results show that -Align provides useful insights about Large Language Model (LLM) development. Across Pythia (1B-12B) checkpoints -Align reveals a non-monotonic three-phase trajectory in representation-head alignment that correlates with downstream capability emergence during pre-training. A similar three-phase trajectory appears in OLMo-2 7B checkpoints, extending the observation beyond Pythia. Across twelve benchmarks, -Align correlates negatively with downstream performance with Spearman correlation -0.923 on Gemma 3 (270M-27B), -0.946 on Qwen2.5 (0.5B-72B) and -0.908 on Qwen3(0.6B-14B). Unlike spectrum-only metrics such as RankMe and -ReQ, -Align maintains a stable direction of association with downstream performance and remains stable across profiling corpora. Finally, in size-matched Qwen2.5 and Qwen3 comparison, pre-training loss favors the weaker model in every pair while width-normalized -Align favors the stronger model in four of five pairs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.