Learning What to Do and How to Do It: Distilling the Decisions Behind Chain-of-Thought Makes Reasoning Robust
Abstract
Distilling multi-step reasoning from large language models into smaller students commonly relies on chain-of-thought (CoT) traces. Yet a CoT step shows how an operation was carried out while leaving the choice of operation implicit, requiring the student to infer what to do from a problem-specific execution. We propose action-execution reasoning, which factorizes each step into a latent *action* that specifies what to do and an *execution* that specifies how to do it. Building on this representation, **ACT**ion–**E**xecution **D**istillation (ACTED) combines supervised learning from verified teacher trajectories with teacher supervision on student rollouts, coupling reliable demonstrations of correct reasoning with corrective guidance at states encountered by the student. Across five reasoning benchmarks, ACTED achieves the highest average Pass@1 and Pass@5 among the compared methods in every tested configuration, distilling Qwen3 teachers of 1.7B–8B parameters into 0.6B-1.7B students. With a 1.7B teacher and 0.6B base student, ACTED improves StrategyQA Pass@5 from 80.4% for the strongest baseline to 89.3%. Diagnostic and controlled perturbation experiments further indicate reduced execution uncertainty and improved robustness to noisy reasoning history.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.