acceptodds
Under review as a conference paper at ICLR 2027

Eigenbases Without Eigenvalues: Exact Amplification Laws for Spectral Preconditioning on Rank-Deficient Gradient Flows

Abstract

Second-order and orthogonal-preconditioned optimizers (Shampoo, SOAP, Muon) deliver strong speedups at pretraining scale yet fail catastrophically in low-rank (LoRA) fine-tuning. This paper's contribution is the exact mechanism of that failure—two amplification laws—not an optimizer. Across five model families (GPT-2, Llama-3.2, Qwen3, Phi-4-mini, Mistral-7B; 124M–7B) and two data domains, the covariances that spectral preconditioners invert are severely rank-deficient: median effective rank 1.1–5.5 against dimensions 768–4096. We prove two exact laws: with a constant eigenvalue floor the amplification equals exactly—independent of learning rate, data, and spectrum—and with a relative ridge it scales as , forcing the ridge to : no interior sweet spot exists. A pre-registered program of 575 runs verifies the laws to seven significant digits and shows free inversion diverges at every learning rate in LoRA and full-parameter training, with the divergence horizon exceeding 30-step calibration probes. The theory identifies the eigenbasis as the only stable channel of spectral information and predicts its own repair: discarding eigenvalue magnitudes while keeping the eigenbasis. The resulting rotation-only variant (SOAP-X)—Muon's principle at pretraining scale, reached here as the theory's prediction for deficient flows—matches or beats AdamW, Muon, and capped-inversion baselines in 6 of 8 configurations, including full-parameter training and a 5000-step horizon; on MMLU-5-shot the diverged optimizer's checkpoint collapses to random-level accuracy (24.89–25.28%) while every stable optimizer finishes 32–43 points above it. A 7B boundary check confirms the cold-start law is dimension-free and finds SOAP-X itself diverging mid-training—a pre-registered miss that three diagnosis blocks then trace to fp32 GPU eigendecomposition returning non-orthogonal eigenvectors on these near-singular covariances; one QR re-orthogonalization per refresh repairs it, verified across seeds and a horizon at 7B. The laws apply to rank-deficient gradient flows (fine-tuning), not pretraining, where covariances fill in and spectral methods succeed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.