Learning Where and How to Train from Model History: Supervised Fine-Tuning with Historical States
Abstract
Supervised fine-tuning (SFT) adapts large language models to downstream tasks through reference responses. We study how a model’s own training history can improve the use of existing supervision by guiding which examples receive training and how their targets shape model updates. We propose WMSS, a framework that pairs the current model with a trainable model initialized from a weaker historical checkpoint. The framework uses current predictive entropy and historical–current entropy differences to allocate training opportunities, and jointly trains the two models through a mixed-logit objective on the selected examples. This produces historical feedback for sample selection and token-level correction while preserving the supervised data pool and original reference targets. We provide a theoretical characterization by deriving an exact decomposition of the mixed residual into residual-scale modulation and candidate-allocation shifts. Experiments across model families demonstrate consistent gains in mathematical reasoning, code generation, and logical reasoning over SFT baselines with the same nominal epoch budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.