Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning
Abstract
A long-standing goal of AI is a model that can continually learn. On heavily post-trained models, direct supervised finetuning (SFT) on new data often forgets and fails to generalize, so it is commonly believed that on-policy training is required. In practice, however, new data to learn from are usually off-policy. On-policy self-distillation (OPSD) tries to bridge this gap by converting off-policy data into on-policy learning signal, but it has been shown to cause reasoning collapse. In this paper, we show that embarrassingly simple off-policy training beats OPSD for continual learning. We show that a major part of SFT's problem lies in its weight damage, and what's worse, the same SFT produces progressively more damage when applied to checkpoints later along the model's training trajectory. To reduce the damage, we show complimentary benefits of (1) scaling the weight update, (2) learning the weight update from an earlier checkpoint (“donor model”) and applying it to the final model, and (3) protecting sensitive weights when needed. Surprisingly, we find that unfinished pretraining checkpoints are the best as the donor model. Across injecting new knowledge, learning from expert traces, and self-improvement, our off-policy training Pareto-dominates SFT and OPSD in injection and retention/generalization while avoiding expensive on-policy sampling. Our work challenges the necessity of on-policy training for continual learning for RL-trained models and points to a better way to continually learn from off-policy data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.