acceptodds
Under review as a conference paper at ICLR 2027

Self-Tidy for Suppressing Self-Resolved Deliberation through Single-Stage Reinforcement Learning

Abstract

Large reasoning models often revisit intermediate conclusions through verification, recomputation, or backtracking, producing unnecessarily long trajectories. Existing length-control methods typically address this inefficiency through response-level constraints, without explicitly accounting for the utility of individual reasoning sentences. To selectively suppress low-value reasoning, we introduce Self-Tidy, a reinforcement-learning method that operates within the standard rollout-to-update process. Self-Tidy scores sentence-level predictive-entropy excursions and subsequent recovery to identify candidate self-resolved deliberation. It withholds positive reinforcement from selected spans while retaining their negative learning signals and preserving the complete sampled trajectories, requiring neither external process supervision nor a separate trace-compression stage. Across three model settings and five mathematical reasoning benchmarks, Self-Tidy reduces mean generated response length by 34.6%–52.2% while maintaining or improving average answer accuracy. These results show that Self-Tidy achieves a competitive balance between answer accuracy and response length without external process supervision or a separate compression pipeline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.