acceptodds
Under review as a conference paper at ICLR 2027

Early-Stopping Frozen Transformer Encoders via Spectral Rank

Abstract

Selecting a transformer backbone for use as a frozen encoder requires committing to depth, head count, width, positional encoding scheme and normalization style. Each combination is conventionally evaluated by training the backbone to convergence and probing it on downstream tasks. Recent theory, validated on synthetic data, predicts a two-phase geometric trajectory for transformers when trained on standard next-token prediction. Transformer representations first compress into low-dimensional, factored subspaces. They then re-expand to capture joint correlations and reduce loss. We investigate a downstream consequence the original theory does not address: whether gains in linear probe accuracy on held-out tasks concentrate near the compression-to-expansion transition, and how these gains depend on the probed target. We test this on two structured-sequence domains: chess and network traffic trace-derived tokens from a specification-compliant Open Radio Access Network (O-RAN) testbed generating data on a variety of telecom equipment combinations, and on natural language using published Pythia training trajectories. We train a family of small next-token language models on these domains, varying architectural choices such as depth, normalization style, RoPE layout and RoPE subspace allocation. Using intra-epoch checkpoints we track both the 95%-variance spectral rank of layer-1 representations and the test accuracy of linear probes trained on held-out tasks whose labels the language model does not see. We show that across the position-aware architectural variants tested, with the no-position control behaving as an expected exception, gains for particular probes concentrate near the point of maximal spectral compression. For these probes, continued language-model training contributes marginal, and in some configurations negative, improvement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.