On the Elicitation and Performance of Compressed Chain-of-Thought Reasoning
Abstract
We study how to compress the reasoning (thinking) that language models produce before they respond, comparing prompting, supervised fine-tuning, and reinforcement learning on reasoning models. For fine-tuning, we train each of three models on telegraphic rewrites of its own reasoning traces and compare it against a control trained on the original traces for the same problems. Fine-tuning reliably elicits compressed reasoning: the share of function words falls from 18–37% to under 2%, and reasoning output becomes 3.5–8.5 shorter. Compressed reasoners are far stronger than truncated reasoning at equal length, but in unconstrained settings they perform worse than the matched control, a gap that 20 more training data narrows but does not close. Prompting with the same specification used to rewrite the source traces does not compress the reasoning of two of the three models, and reinforcement learning with a function-word reward shortens reasoning without changing its style. Overall, we find that compressed reasoning is useful only at very small token budgets; otherwise, disabling reasoning achieves higher accuracy while still using far fewer tokens than full reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.