acceptodds
Under review as a conference paper at ICLR 2027

When Does Reasoning SFT Generalize Better Than GRPO?

Abstract

Supervised fine-tuning (SFT) is often associated with memorization, whereas reinforcement learning (RL) is credited with stronger generalization in large language models (LLMs). We show that sufficiently trained SFT can achieve stronger out-of-distribution (OOD) generalization than Group Relative Policy Optimization (GRPO). Tracking SFT with long chain-of-thought (Long-CoT) supervision and GRPO over multiple epochs under matched training-problem sequences and checkpoints, we find that SFT surpasses GRPO across consecutive checkpoints after sufficient training, whereas GRPO is generally stronger or comparable at early checkpoints. Response analysis shows that strict verification language, identified by markers that directly check intermediate or final results, remains common in SFT OOD generations. The SFT–GRPO difference in verification density is already visible near GRPO’s early high-performing checkpoint, before its later weakening in our matched run. Crucially, SFT’s OOD advantage is not explained by simply generating longer reasoning: successful SFT responses are typically shorter and use verification more densely than failed responses. With further training, SFT’s OOD performance declines from the stronger levels it reaches, although GRPO does not overtake SFT again within the observed training horizon. These results challenge the view that SFT memorizes while RL generalizes. SFT’s generalization gains can be missed by stopping too early and eroded by training too long.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.