TokenDistill: Compressing CoT via Iterative Token Distillation with Continuous Prompt Transformation
Abstract
Chain-of-Thought (CoT) prompting can substantially improve the final-answer accuracy of large language models (LLMs) by enabling step-by-step reasoning. However, the intermediate tokens it generates incur additional computational and memory overhead, thereby increasing deployment costs. Existing CoT compression methods either attain limited compression ratios, rely on auxiliary external models, or require extra comparison samples during optimization, which introduces considerable training overhead. To address these limitations, we propose TokenDistill, a novel training pipeline that enables the model itself to iteratively distill CoT tokens in an adaptive manner, achieving aggressive CoT compression while largely preserving final-answer accuracy. At the core of TokenDistill is an iterative prompt transformation mechanism, which enables the model to internalize the requirement for concise outputs from explicit prompting into its own generation behavior. Specifically, we first augment the original input prompts with a concise-output instruction to obtain concise outputs. We then fine-tune the model using the original inputs without the concise-output instruction and the generated concise outputs as supervised input-output pairs, thereby internalizing the concise-output behavior into the model. Furthermore, we repeat the aforementioned procedure over multiple rounds, enabling the model to progressively push its compression limits beyond the previously attained level. Extensive experiments demonstrate that TokenDistill consistently outperforms state-of-the-art methods in CoT compression, incurring only minimal accuracy loss and in some cases even slight accuracy gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.