acceptodds
Under review as a conference paper at ICLR 2027

Not All Short Chains Are Equal: How CoT Compression Affects SFT and RL

Abstract

Chain-of-thought (CoT) compression has attracted growing interest as a way to reduce training and inference costs. However, the compression ratio and accuracy may not fully reflect the value of compressed trajectories as training data. Similar compression ratios may lead to different outcomes in supervised fine-tuning (SFT) and subsequent reinforcement learning (RL). To investigate this, we conduct an end-to-end study on mathematical reasoning, comparing uncompressed data, prompt-based Chain-of-Draft compression, and self-distillation-based CRISP compression. On Qwen3-8B, SFT on Chain-of-Draft data yields 12.78 percentage points higher AIME accuracy than SFT on CRISP data. Its AIME advantage is concentrated on problems that the starting model can solve occasionally but inconsistently. In subsequent RL, Chain-of-Draft retains higher final accuracy than CRISP. These findings suggest that the value of short-CoT training data depends not only on response length, but also on how well it preserves the model's ability to solve problems it initially solves inconsistently. Since SFT also provides the initialization for RL, evaluating this training value should consider both immediate SFT outcomes and final performance after RL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.