Rethinking Supervised Fine-Tuning through Policy Projection: From Uniform Imitation to Selective Adaptation
Abstract
Supervised fine-tuning (SFT) has become a key recipe for adapting large language models (LLMs) to different domains through imitation of task-specific demonstrations. However, its deterministic objective applies indiscriminate learning to each reference token, which may lead to overfitting and ignore whether the supervision is compatible with the model's current predictive state. We first develop a structured policy projection framework that views fine-tuning as controlled movement in policy space and provides a principled design space by characterizing each token-level projection in terms of where the policy is directed with minimum shift and how strongly it contributes to optimization. This formulation reveals standard SFT as a rigid case with a degenerate projection target and uniform strength. Based on this, we further propose Modulated Fine-Tuning (MFT), which combines a compatibility-adaptive relaxed projection with an information-aware projection strength to modulate token-wise supervision according to the current policy, consolidating compatible and informative supervision while applying more conservative updates to strongly conflicting targets. Experiments across diverse domains show that MFT consistently outperforms SFT, improving average performance by +6.96 on mathematical reasoning and +10.8 on general domain tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.