OPR: On-Policy Replay for Continual Learning in Large Language Models
Abstract
Large language models can continually learn new tasks through sequential supervised fine-tuning (SFT), but remain susceptible to catastrophic forgetting. We propose On-Policy Replay (OPR), a generative replay method that uses selected responses from the current model as supervision for previously learned tasks. After each task, the current checkpoint generates responses to historical prompts. High-scoring prompt–response pairs are then mixed with the next task’s labeled data for SFT using the standard cross-entropy objective. We introduce two response selection strategies: OPR-RU uses task rewards, whereas OPR-SC uses model confidence without requiring historical labels for replay construction. By regenerating replay responses at each task boundary, OPR updates the supervision for past tasks to reflect the model’s current output distribution. On the TRACE benchmark, OPR-RU achieves higher average task accuracy and lower forgetting than LAMOL and SDFT at the same 1% replay budget on both Qwen3-4B and Qwen3-8B. Increasing the replay budget to 10% yields a forgetting rate of 0.19% on Qwen3-4B. These results demonstrate that OPR preserves previously learned task capabilities, providing an effective approach to continual learning in large language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.