Self-Tidy for Suppressing Self-Resolved Deliberation through Single-Stage Reinforcement Learning
Abstract
Large reasoning models often revisit intermediate conclusions through verification, recomputation, or backtracking, producing unnecessarily long trajectories. Existing length-control methods typically address this inefficiency through response-level constraints, without explicitly accounting for the utility of individual reasoning sentences. To selectively suppress low-value reasoning, we introduce Self-Tidy, a reinforcement-learning method that operates within the standard rollout-to-update process. Self-Tidy scores sentence-level predictive-entropy excursions and subsequent recovery to identify candidate self-resolved deliberation. It withholds positive reinforcement from selected spans while retaining their negative learning signals and preserving the complete sampled trajectories, requiring neither external process supervision nor a separate trace-compression stage. Across three model settings and five mathematical reasoning benchmarks, Self-Tidy reduces mean generated response length by 34.6%–52.2% while maintaining or improving average answer accuracy. These results show that Self-Tidy achieves a competitive balance between answer accuracy and response length without external process supervision or a separate compression pipeline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.