Training Schedule Order Selects Isospectral Matrix Completions
Abstract
Training schedules can change predictions even when they use the same learning rates and noise levels equally often. We construct a rank-two Gram completion problem in which reversing a control cycle changes unobserved predictions while preserving training fit and the entire matrix spectrum after the same final noiseless gradient flow. The mechanism is covariance lag: fluctuations retain earlier controls, and Stokes' theorem expresses the slow-cycle response as curvature flux through the control loop. By canceling the exact drift produced by fixed controls, we prove persistent separation over many cycles, vanishing fluctuations and quantitative bounds for composing schedules, and certify the separation at the experimental noise level. We turn this response into Checkpoint Response (CR), a method for forecasting and directing schedule effects, and test it on matrix completion and billion-parameter neural models. Across regression families, CR reduces median forecast error from 7.3% to 11.7%, and in native classification, it reduces full-output forecast error by 49-53%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.