acceptodds
Under review as a conference paper at ICLR 2027

Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training

Abstract

Pre-training large language models is dominated by the memory cost of storing full-rank weights, gradients, and optimizer states. Low-rank pre-training has emerged to address this, and the space of methods has grown rapidly. A central question remains open: do low-rank methods produce models that generalize comparably to full-rank training, or does the rank constraint fundamentally alter the solutions reached? Existing comparisons rely almost entirely on validation perplexity from single-seed runs, often carried forward from prior literature. Yet perplexity is a poor proxy for solution quality — two methods can match on perplexity while converging to very different loss landscape regions and internal representations. To look beyond perplexity, we introduce BASIN (Barrier, Activations, Spectra, and Intrinsic geometry of Networks), a framework that characterizes the solutions found by five low-rank pre-training methods — GaLore and Fira (memory-efficient optimizers), CoLA and SLTrain (architecture reparameterizations), and ReLoRA (adapter-style updates with periodic resets) — against full-rank training at three model scales (60M, 130M, 350M) under three seeds. We evaluate each along 16 metrics across four dimensions: 1-D loss landscape along random/top-K PCA directions, 1-D interpolation between checkpoints, spectral structure of the weights and learned updates, and activation similarity to full-rank training. We also examine how each pre-trained model adapts after task and instruction fine-tuning, and continual pre-training with and without replay. We show that low-rank methods are not equivalent to full-rank training, nor to one another, even when their validation perplexity is close. Each method converges to a distinct basin with different sharpness and spectrum. Low-rank method activations drift further from full-rank in later layers and as training progresses. These differences shape post-training behavior. For example, the path from the pre-trained to the fine-tuned model crosses a barrier for ReLoRA and climbs a steep wall for SLTrain, matching their high forgetting. Finally, in two case studies, BASIN traces how Fira and SLTrain deviate from full-rank, and a simple fix guided by this diagnosis brings each method closer to full-rank training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.