acceptodds
Under review as a conference paper at ICLR 2027

Should SFT Prefer the Familiar? Diversity Collapse and Recovery via Model Merging

Abstract

Supervised fine-tuning (SFT) is widely used for downstream adaptation and often serves as initialization for further reinforcement learning (RL). Recently, several methods have improved SFT by upweighting from tokens that are already familiar to the current policy. Some approaches such as DFT prioritize high-probability tokens within external data, while some approaches train on self-generated trajectories, which naturally contain tokens that are familiar to the current model. Although these methods have good post-SFT performance, we empirically find that preferring familiar tokens in SFT can substantially reduce policy diversity and lead to worse subsequent RL performance. We further show theoretically that such methods create a self-reinforcing dynamic that progressively contracts the model's effective support, causing the model to become overly concentrated and reduce the diversity. To mitigate this effect, we propose a simple recovery approach: merging the fine-tuned model with its pretrained base model before RL. This approach can restore diversity lost during fine-tuning while preserving the gains in post-SFT performance. Across multiple mathematical reasoning and code generation benchmarks, base-model merging consistently outperforms recent SFT variants and standard regularization approaches such as KL-divergence regularization, yielding improvements in both post-SFT accuracy and subsequent RL performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.