A Unifying Lens on Fine-Tuning Through Target Distribution Design
Abstract
Supervised fine-tuning (SFT) typically maximizes the likelihood of every token in a demonstrated trajectory. However, an observed token can be non-unique, noisy, or misaligned with the model prior. Strictly fitting toward this one-hot target may be suboptimal, especially when the pretrained model encodes a rich knowledge prior. In this work, we reinterpret SFT as target distribution design: instead of studying only the loss objective, we analyze the token-level target that a loss drives the model to match. We introduce the Target framework, which decomposes SFT supervision into two explicit choices: (1) how strongly to rely on the observed token, and (2) how to allocate the remaining probability mass over alternatives. This perspective unifies many existing SFT variants as implicit choices of the target distribution. Building on this view, we propose Target-SFT which directly constructs the desired target by adaptively balancing label imitation with a teacher-guided residual distribution. Across three datasets and seven models spanning math and medical reasoning, this method consistently outperforms all baselines, showing the effectiveness of this target-based approach. Overall, our formulation reveals a fundamental design perspective and opens a broader search space for training objectives.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.