Learning from How Large Language Models Have Changed
Abstract
Supervised fine-tuning (SFT) and on-policy distillation (OPD) improve language models through demonstrations and teacher feedback. However, the learner’s current state alone may not reveal how these signals should be allocated to support further task improvement. We introduce REWIND, which uses the model’s responses to prior training to guide fine-grained allocation of training signals within each sequence while preserving each supervision term’s learning direction. The same approach applies to both SFT and OPD without requiring an auxiliary feedback-generation loop. We provide theoretical justification by showing that REWIND can locally reduce task risk at a population SFT stationary point and by establishing teacher consistency for fixed-prefix OPD. Experiments on six mathematical reasoning benchmarks at two model scales demonstrate consistent performance gains in both SFT and OPD. Matched ablations further support the importance of token-specific signal allocation, the training-origin reference, and information from competing tokens.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.