Tether: Decision Retention for Continual Fine-Tuning of LLMs
Abstract
Continual fine-tuning can cause catastrophic forgetting in language models. Changes in old-task accuracy reflect a net effect: repaired errors offset lost correct predictions. We instead measure and constrain decision drift, the fraction of old-task inputs whose decisions differ from those of the model at the end of that task, the teacher. Drift bounds forgetting, both lost correct answers and the loss in accuracy, whereas accuracy does not bound drift; it needs no labels and has an exact finite-sample upper bound on held-out inputs. We propose Tether, a framework that stores real or synthetic inputs with the model's outputs as records and regularises later training with any continuous loss that upper-bounds record drift. For forward-KL distribution matching on records, we derive the exact per-record critical radius below which the recorded decision is preserved; it shows that standard KL also penalises distributional changes that leave the recorded decision intact. Two new instances, Tether-Margin and thresholded KL (Tether-TKL), are zero on a region around the teacher and skip satisfied records before back-propagation. On the O-LoRA benchmark with T5-large and Llama-2-7B-chat and on six generation tasks with Llama-3-8B-Instruct, without retaining per-example ground-truth labels, the three instances reach final average scores comparable to DER++'s, from 1.1 points below to 1.5 above. Tether-TKL and KL change fewer earlier decisions than DER++, KL losing about half as many correct answers, and on generation all three keep 50–79% of earlier free-form outputs unchanged, against 38% for DER++. On the O-LoRA benchmark, with fewer retained records in aggregate than the 2% replay and DER++ buffers, all three also preserve more of the model's earlier behaviour than labelled replay, changing fewer decisions in every setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.