Stopping Soft Reasoning before It Loops: From Impasse-Oriented Detection to Event-Level Intervention
Abstract
Large reasoning models (LRMs) achieve strong performance by generating long chains of thought (CoT), but discrete token sampling discards distributional information at each step. Training-free soft reasoning retains this information through probability-weighted embeddings, yet repeated (soft) feedback can prevent termination. We show that this failure is not merely out-of-distribution drift by soft representations: soft trajectories can amplify an impasse-induced recurrent regime already present in hard decoding, in which non-adjacent semantic events revisit similar hidden-state regions and predictive entropy progressively contracts. Motivated by this finding, we propose E-TRACE, a training-free, event-level detector that identifies persistent non-local recurrence in projected hidden-state distributions. E-TRACE stores entropy-decreasing events as rollback candidates and, once recurrence is confirmed, restores the nearest candidate outside the recurrent region, truncates the recurrent suffix, and switches to hard decoding. Experiments show that the method can accurately predict non-termination and improve overall reasoning performance, achieving higher answer accuracy with substantially fewer generated tokens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.