Learning Dynamics-Informed Learning Curve Extrapolation
Abstract
As the computational demands of machine learning training keep growing, how to obtain more high-quality models under limited resources has attracted increasing attention. Recently, Computational Resource Efficient Learning (CoRE-Learning) has been proposed to address this problem. A key component of CoRE-Learning is learning curve extrapolation (LCE), which predicts future model performance from partially observed training trajectories to guide resource allocation. Existing LCE methods largely rely on curve families or empirical descriptors, leaving the underlying learning process implicit. Nonetheless, we observe that similar early curves can yield different future improvements under the same learning-rate schedule, while a shared checkpoint can evolve differently under different schedules. These observations motivate us to construct learning dynamics-informed generative priors for LCE. We propose Learning Dynamics-Informed PFN (LD-PFN), which builds a generative prior from a reduced spectral model of learning dynamics. A spectral surrogate represents the remaining loss as components with different decay rates, and learning-rate schedules map resource usage to effective optimization progress. We train a PFN on trajectories sampled from this prior for probabilistic extrapolation of validation cross-entropy, without task-specific updates or access to model internals at prediction time. Experiments across diverse fine-tuning scenarios and evaluation settings show that LD-PFN improves predictive accuracy and probabilistic quality over existing baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.