LMATE: Learning New Tasks without Forgetting through Teacher-Guided Revision
Abstract
Continual instruction tuning seeks to fine-tune an instruction tuned model to a new target task while avoiding to lose its general capabilities. Supervised Fine-Tuning (SFT) trains the model to mimic ground-truth demonstrations from the dataset that may differ substantially from their existing response distribution, causing unnecessary changes and leading to catastrophic forgetting. We present LLM as a Teacher (LMATE), which adopts a teacher/student approach that constructs task-informed training targets from the student’s own answers. For each prompt, the student generates an initial response, and a stronger teacher with access to the ground-truth demonstration provides targeted feedback. Conditioned on its initial response and this feedback, the student produces a revised response through in-context learning. We then fine-tune the student on these revised answers. To limit learning from suffixes that may drift after an early correction, we apply the training loss only to the revised prefix ending at the -th token-level difference. This concentrates the update on the earliest teacher-guided corrections while excluding more extensively revised suffixes. We evaluate LMATE on scientific question answering, tool use, and factual knowledge acquisition, measuring retention with IFEval and MMLU-Pro. LMATE achieves a competitive adaptation–retention trade-off across these tasks. On Tool-Use it reaches 66.19% accuracy compared with 61.03% for SFT, and for Timely Events it reaches 22.44 F1-score compared with 17.19 for SFT, while retaining higher IFEval and MMLU-Pro accuracy on both datasets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.