Free to Ignore: Cornerman Decoding for Distilled Reasoners
Abstract
Distilling long chain-of-thought (CoT) traces from stronger teachers offers a practical route to improving reasoning in small language models. These traces demonstrate both solution content and the organization of reasoning. Yet our analysis shows that students can produce useful intermediate reasoning without using it effectively to advance their solutions. This motivates guidance on how students develop, revise, and complete solutions during generation. We introduce Cornerman Decoding, which pairs a frozen student with a learned local adviser. The adviser observes only the latest student paragraph and either remains silent or suggests continuing, switching approaches, checking a claim, or finishing, without supplying new solution steps or answers. The student retains the full problem and reasoning history and is free to follow or ignore the advice. Across reasoning benchmarks, we evaluate the same frozen 1.5B adviser with 1.5B and 7B students fine-tuned on complete CoT traces. Compared with decoding the same students without advice, Cornerman improves macro-average accuracy by 5.5 and 1.0 percentage points, respectively, while reducing mean student-generated tokens by 45.2% and 21.1%. Paired case studies further illustrate how guidance can help students reach correct answers that do not appear in their uncoached trajectories.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.