Forecasting Post-Checkpoint Training with Moving-Frame Dynamics
Abstract
We study how far the future trajectory of a training run can be forecast from its current checkpoint. Because learned features continue to evolve, coordinates fixed at the checkpoint become progressively inaccurate. We introduce a moving-frame predictor that transports the checkpoint-conditioned mean and covariance with the evolving representation, using optimizer-induced feature motion to update the frame. This models the forecastable component of the continuation. We then establish a complementary, method-independent limit: the irreducible state-prediction MSE induced by post-checkpoint stochasticity is locally linear in the forecast horizon, whereas the additional error from freezing feature evolution admits a local quartic bound. The resulting error budget explains how inaccuracies accumulate over a completed forecast. Experiments on end-to-end multilayer perceptrons and Transformers for next-token prediction on real text verify the predicted error structure and show that the moving-frame predictor substantially outperforms checkpoint-fixed and learning-curve extrapolation baselines. Together, these results separate forecast error that can be reduced by modeling representation dynamics from uncertainty intrinsic to the future training path.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.