Should SFT Prefer the Familiar? Diversity Collapse and Recovery via Model Merging
Abstract
Supervised fine-tuning (SFT) is widely used for downstream adaptation and often serves as initialization for further reinforcement learning (RL). Recently, several methods have improved SFT by upweighting from tokens that are already familiar to the current policy. Some approaches such as DFT prioritize high-probability tokens within external data, while some approaches train on self-generated trajectories, which naturally contain tokens that are familiar to the current model. Although these methods have good post-SFT performance, we empirically find that preferring familiar tokens in SFT can substantially reduce policy diversity and lead to worse subsequent RL performance. We further show theoretically that such methods create a self-reinforcing dynamic that progressively contracts the model's effective support, causing the model to become overly concentrated and reduce the diversity. To mitigate this effect, we propose a simple recovery approach: merging the fine-tuned model with its pretrained base model before RL. This approach can restore diversity lost during fine-tuning while preserving the gains in post-SFT performance. Across multiple mathematical reasoning and code generation benchmarks, base-model merging consistently outperforms recent SFT variants and standard regularization approaches such as KL-divergence regularization, yielding improvements in both post-SFT accuracy and subsequent RL performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.