acceptodds
Under review as a conference paper at ICLR 2027

Learning What to Do and How to Do It: Distilling the Decisions Behind Chain-of-Thought Makes Reasoning Robust

Abstract

Distilling multi-step reasoning from large language models into smaller students commonly relies on chain-of-thought (CoT) traces. Yet a CoT step shows how an operation was carried out while leaving the choice of operation implicit, requiring the student to infer what to do from a problem-specific execution. We propose action-execution reasoning, which factorizes each step into a latent *action* that specifies what to do and an *execution* that specifies how to do it. Building on this representation, **ACT**ion–**E**xecution **D**istillation (ACTED) combines supervised learning from verified teacher trajectories with teacher supervision on student rollouts, coupling reliable demonstrations of correct reasoning with corrective guidance at states encountered by the student. Across five reasoning benchmarks, ACTED achieves the highest average Pass@1 and Pass@5 among the compared methods in every tested configuration, distilling Qwen3 teachers of 1.7B–8B parameters into 0.6B-1.7B students. With a 1.7B teacher and 0.6B base student, ACTED improves StrategyQA Pass@5 from 80.4% for the strongest baseline to 89.3%. Diagnostic and controlled perturbation experiments further indicate reduced execution uncertainty and improved robustness to noisy reasoning history.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.