acceptodds
Under review as a conference paper at ICLR 2027

Spectral Monitors Detect Change, Not Damage, in LLM Post-Training

Abstract

Can hidden-state spectra distinguish harmful post-training updates from benign adaptation? A fixed-readout construction shows that identical spectra can coexist with arbitrarily different predictive losses—an identifiability limit, not an impossibility theorem for empirical alarms. We test the empirical question with pre-registered shared-prefix stress tests on three open models, run on a corrected training/probe split with fresh seeds. Extreme duplication damages held-out NLL by to nats while raw RankMe rises in 9/9 model–seeds, and the same high-LR damage moves RankMe down on OLMo-2-1B but up on Qwen3-1.7B: the sign of a spectral response tracks the (regime, architecture, layer) tuple, not damage. A trajectory-level split-conformal protocol then compares spectral and loss monitors at an equal false-alarm budget on the development model (Qwen3-0.6B): on abrupt damage every spectral family alarms at the same measurement as a plain loss detector; on a gradual learning-rate ramp spectra alarm 80–110 steps earlier, but CUSUM and Page–Hinkley detectors on the same loss stream recover about half (CUSUM) and a quarter (Page–Hinkley) of that lead, and all of these detectors also fire on a non-damaging duplication ramp and on a benign control that improves its target ( nats) with no significant task-accuracy change. Only a loss monitor with a non-degenerate threshold separates damage from benign change; a pre-registered matched-threshold control shows the sidedness of the detector is immaterial. An earlier campaign affected by a training/probe exposure defect is disclosed, repaired, and reported side by side; its structure replicates cleanly. In these settings, the tested spectral diagnostics measure change, not damage.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.