acceptodds
Under review as a conference paper at ICLR 2027

Optimizer Lineage: The Pretraining Optimizer Leaves a Causal Spectral Fingerprint on Base-Model Weights

Abstract

A base model is usually treated as a black box that downstream users prune, quantize, LoRA-adapt, or merge—as if these operations depend only on its architecture and final weights. We show that its weight spectrum also carries a durable, causal imprint of a choice fixed long before any user touches the checkpoint: the optimizer it was pretrained with. Adaptive optimizers such as Adam apply per-coordinate, effectively low-rank updates, whereas spectral optimizers such as Muon orthogonalize each update via a matrix-sign (msign) step, spreading energy across the full singular spectrum. Our single load-bearing result is a causal spectral fingerprint (H1). Under a clean-room control that pretrains 124M GPT pairs varying only the optimizer—architecture, data, initialization seed, and step count all held fixed— Muon raises the per-layer stable rank of hidden matrices by 6–14× (from  17 to  121) and lightens their spectral tail relative to Adam. The effect is unusually robust: it holds consistently across every layer, replicates across two independent seeds, and reproduces at two scales (124M and 350M). It is causal rather than a loss artifact—under an iso-loss control that trains Adam to matched (or lower) loss the gap persists and grows, at both scales (a 350M pair keeps a  12–25× stable-rank gap, p=7e-15)—so it tracks update geometry, not training progress. The fingerprint has a concrete deployment consequence: it sets a checkpoint’s compression budget. Muon-lineage weights are 5–7× less low-rank-truncatable and  29× less amenable to 4-bit quantization than Adam-lineage weights (again robust to the iso-loss control), so the pretraining optimizer—not just the architecture and final weights—determines how cheaply a base model can be compressed. Optimizer lineage is thus a measurable, causal, and deployment-relevant property that deserves to travel with a checkpoint as first-class metadata.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.