acceptodds
Under review as a conference paper at ICLR 2027

MENTOR-SDFT: Making Explicit Knowledge Transferable via On-policy References for Continual Self-Distillation Fine-Tuning

Abstract

Continual fine-tuning requires large language models (LLMs) to acquire new capabilities while retaining existing knowledge. Self-Distillation Fine-Tuning (SDFT) offers a practical approach to this problem by using a reference-conditioned self-teacher to provide on-policy soft targets. However, we find that its supervision is often weak: reference conditioning alone does not guarantee that the self-teacher can transform the reference into transferable supervision, thereby undermining the plasticity–stability trade-off in continual learning. To address this limitation, we propose MENTOR-SDFT (Making Explicit kNowledge Transferable via On-policy References for SDFT), a framework that strengthens self-teacher guidance by improving both supervision construction and utilization. MENTOR-SDFT introduces Agent-Guided Prompt Evolution to optimize teacher conditioning and Confidence-guided Soft Trajectory Distillation to adaptively regulate supervision according to teacher confidence and trajectory position. Experiments across continual-learning sequences and multiple LLM backbones show that MENTOR-SDFT improves final average accuracy and reduces forgetting compared with sequential SDFT baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.