acceptodds
Under review as a conference paper at ICLR 2027

AID: Asymmetric Intermediate-Answer Distillation for Budget-Constrained Long-CoT Inference

Abstract

Supervised fine-tuning (SFT) on long chain-of-thought (long-CoT) traces is a common way to transfer reasoning ability from a large teacher to a smaller student, but this transfer can be brittle when the inference-time reasoning budget is limited. The distilled students often fail to finish reasoning within the available budget, and forcing an answer from an unfinished state does not reliably recover the correct solution. We introduce AID (Asymmetric Intermediate-Answer Distillation), which retains supervision on the full teacher trace while additionally training the student to predict the final answer from randomly sampled intermediate states. AID uses an asymmetric attention mask to keep the auxiliary answer branch out of the reasoning stream's forward context, allowing full-CoT imitation and intermediate-state answerability to be learned jointly. AID requires neither additional data generation nor extra parameters, and leaves inference unchanged. Experiments with Qwen3-0.6B trained on OpenR1-Math and KodCode show that plain SFT degrades sharply on splits that combine long traces with harder problems, whereas AID substantially mitigates this degradation. Under long-trace distillation, accuracy improves from 22.9% to 37.0% on MATH-500 and from 25.6% to 35.6% on MBPP+. In a truncation-based evaluation on held-out validation problems, which reveals progressively longer prefixes of a reference reasoning trace, the normalized area under the accuracy curve rises from 0.191 to 0.419 on math and from 0.321 to 0.450 on code. Our results indicate that intermediate-state answerability provides a complementary training signal that makes long-CoT distillation more robust under constrained inference budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.