Not All Short Chains Are Equal: How CoT Compression Affects SFT and RL
Abstract
Chain-of-thought (CoT) compression has attracted growing interest as a way to reduce training and inference costs. However, the compression ratio and accuracy may not fully reflect the value of compressed trajectories as training data. Similar compression ratios may lead to different outcomes in supervised fine-tuning (SFT) and subsequent reinforcement learning (RL). To investigate this, we conduct an end-to-end study on mathematical reasoning, comparing uncompressed data, prompt-based Chain-of-Draft compression, and self-distillation-based CRISP compression. On Qwen3-8B, SFT on Chain-of-Draft data yields 12.78 percentage points higher AIME accuracy than SFT on CRISP data. Its AIME advantage is concentrated on problems that the starting model can solve occasionally but inconsistently. In subsequent RL, Chain-of-Draft retains higher final accuracy than CRISP. These findings suggest that the value of short-CoT training data depends not only on response length, but also on how well it preserves the model's ability to solve problems it initially solves inconsistently. Since SFT also provides the initialization for RL, evaluating this training value should consider both immediate SFT outcomes and final performance after RL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.