In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners
Abstract
Stronger learning signals do not necessarily produce a stronger small reasoning model. Off-policy supervised fine-tuning (SFT) supplies stronger traces that may be incompatible with the student. On-policy post-training instead remains constrained by the quality of the student's own trajectories: reinforcement learning (RL) provides little learning signal because a small student rarely produces a correct rollout. Consequently, effective post-training must supply stronger reasoning in a form the student can learn from. We hypothesize that a small number of tokens with very low probability under the student can make an otherwise strong reasoning trace difficult to learn. To test this hypothesis, we introduce Interleaved-Policy Distillation (IPD), which generates SFT data balancing teacher guidance and student compatibility. During generation, the student replaces teacher proposals that have low probability under its own policy, allowing the teacher to continue reasoning from a prefix the student can support. IPD improves reasoning across model families, datasets, and varying degrees of student intervention; for example, it raises the seven-benchmark average from 24.55 to 29.72 on Qwen3-0.6B/s1K. It outperforms baselines with comparable training budgets and broadens problem-solving coverage, while ablations show that filtering traces or masking losses cannot match the gains from revising trajectories during generation. Analyses link transfer failures to the rare tokens the student finds least likely and show that the source of the tokens also matters. Small models reason better when they learn strong solutions in a form they can reach in their own words.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.