Logic Density Pruning: Mitigating Overthinking via Step-Level Chain of Thought Compression
Abstract
Large reasoning models often improve on complex tasks by generating longer Chains-of-Thought (CoT). However, prolonged reasoning trajectories do not strictly guarantee improved utility. The phenomenon of generating redundant paraphrasing, loose explanations, and unproductive branches, which substantially inflates token usage and latency, is often termed *overthinking*. Prior work attempts to control reasoning length via prompt engineering, teacher-driven rewriting and distillation, or training-based constraints, but it commonly lacks an interpretable metric over reasoning steps that attributes each reasoning step to the final answer, making it difficult to reliably balance compression and fidelity. Compared to these methods, this paper proposes Logic Density Pruning (LDP), a CoT compression and training framework. Rather than relying on token-level importance, LDP computes step importance based on a unified criterion, *logic density*, which is constructed from two complementary scores at the step level—Mutual Information (MI) and attention citation. Based on this logic density, LDP performs reasonable CoT compression. To reduce the risk of mistakenly removing truly critical steps, this paper introduces a step perturbation check that verifies whether removing a candidate step induces a significant shift in the model’s subsequent generation distribution, and selectively recovers such steps. Notably, when the compressed CoTs are applied to Qwen3-8B, reasoning tokens are reduced by 80.2% (from 295.5 to 58.6) on CommonsenseQA, with only a 1.9% performance drop being observed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.