acceptodds
Under review as a conference paper at ICLR 2027

Distilling Model-Inherent Reasoning Circuit for Large Language Models

Abstract

Knowledge Distillation (KD) is widely adopted to compress Large Language Models (LLMs) by aligning the output distributions of teacher model and student model. However, output matching alone does not explicitly characterize the internal computations reflecting the teacher's reasoning capability, and directly aligning the representations of paired Transformer components risks disrupting the student's intrinsic reasoning pattern. To address these limitations, we propose model-inherent Reasoning Circuit Distillation (RCD). RCD characterizes model-inherent reasoning patterns as critical circuits formed by Transformer edges, with a separate circuit identified for each model using Edge Attribution Patching with Integrated Gradients. Subsequently, to transfer the teacher's reasoning capability, RCD aligns the predictive effects of both circuits by matching the logit variations induced by circuit attenuation. Theoretically, we establish guarantees for recovering the ground-truth reasoning circuits and show that the empirical RCD loss controls the population discrepancy in predictive effects between the ground-truth circuits of the teacher model and the student model. Experiments on mathematical reasoning and code generation show that RCD consistently outperforms state-of-the-art KD methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.