Not Every Success Is Worth Imitating: Rethinking Data Selection for Reasoning
Abstract
Rejection-sampling supervised fine-tuning (RS-SFT) improves a language model by training on its own generated responses that reach the correct answer. Despite its simplicity and low cost, RS-SFT is commonly regarded as a weaker alternative to reinforcement learning (RL) for reasoning. We identify an overlooked limitation in its data selection: a response that eventually succeeds is not necessarily a desirable trajectory to imitate. Standard outcome-based selection checks only the final answer and can therefore admit responses that make failed attempts before reaching the correct answer or continue generating meaningless text after already answering. Under iterative self-training, these undesirable behaviors can be learned, regenerated, and selected again; in our experiments, post-answer degeneration among selected responses increases from 5.2% to 34.7%. Motivated by this observation, we propose Imitation-Aware Trajectory Selection (IMIT), a lightweight selection method that retains a correct response only if it contains exactly one final answer and terminates immediately after it. IMIT changes neither the training objective nor the training budget. On Qwen3-1.7B-Base and MATH-500, IMIT improves first-answer pass@1 over standard RS-SFT by 2.12 percentage points in one-round training and by 4.37 points under iterative self-training. Controlled experiments show that these gains arise primarily from which responses are selected, rather than from the reduced number of training problems. Despite its more selective training data, IMIT preserves solution diversity close to that of the base model and reaches 89.0% first-answer pass@32, compared with 86.4% for GRPO, although GRPO achieves substantially higher single-sample accuracy. Our results show that outcome correctness alone is insufficient for selecting self-generated responses for imitation, and suggest that the selection criterion is a critical component of effective self-training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.