acceptodds
Under review as a conference paper at ICLR 2027

TailSFT: Filtered Fine-Tuning Improves Post-Training Performance

Abstract

Reinforcement learning post-training drives reasoning and agentic capabilities in modern AI systems, yet a growing body of work shows that it is most effective when used to fine-tune an already capable base model. We question whether existing pipelines yield models that are most suitable for reinforcement learning. Building on prior work highlighting the role of coverage and pass@K as predictors of post-RL performance, we design a simple modification to supervised fine-tuning, TailSFT, which filters out already fit sequences during training, thereby focusing learning on under-modeled regions, or the tail, of the data distribution. We justify and validate the design choices in TailSFT, particularly the specific filtering criteria, through a combination of controlled experiments and theoretical analysis. On OLMo-3 7B, TailSFT improves pass@16 in 15 of 18 math and coding comparisons, with gains up to 16.8 percentage points, while incurring minimal computational overhead. Initializing GRPO from TailSFT checkpoints yields up to 3.9 percentage points of pass@1 improvement in the reported comparisons, consistent with the hypothesis that broader coverage supports subsequent RL. We further introduce a lightweight diagnostic for identifying settings where TailSFT is most likely to help. More broadly, our results motivate a principled, stage-aware approach to model development, in which intermediate checkpoints are judged by how effectively they support subsequent training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.