Trust-region Constraint Improves Continual Learning of Self-distillation Fine-tuning
Abstract
Continually adapting large language models to new domains risks erasing capabilities acquired during pretraining. Self-Distillation Fine-Tuning (SDFT) mitigates this problem by using a demonstration-conditioned version of the model as an on-policy teacher. We show, however, that although this teacher is substantially closer to the student than one-hot supervision on average, its teacher–policy divergence remains strongly heavy-tailed across model families and adaptation domains. A controlled down-weighting intervention further shows that high-divergence teachers contribute disproportionately to forgetting. Motivated by this finding, we introduce Trust-Region Constrained Self-Distillation Fine-Tuning (TRSD), which constructs a sample-wise trust-region-constrained teacher distribution that remains faithful to the demonstration-conditioned teacher while limiting its divergence from the current student. Formulated over the -divergence family, TRSD induces distinct teacher-distribution geometries for forward KL, reverse KL, and squared Hellinger distance without requiring an additional model forward pass beyond SDFT. Across tool use, scientific question answering, and medical reasoning, TRSD improves instruction-following over SDFT, generally without requiring weaker task accuracy. These gains persist under sequential adaptation and broad hyperparameter sweeps. Our results establish teacher calibration as an important principle for continual self-distillation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.