ReTrail: Hesitation-Guided Backtracking for Test-Time Reasoning
Abstract
Large reasoning models (LRMs) often revise earlier steps through reflection, but insufficient reflection leaves mistakes uncorrected, whereas excessive reflection leads to repetitive reasoning. Existing approaches that regulate reflection using fixed schedules cannot adequately address this difficulty because they do not adapt to whether the model needs to reconsider its reasoning or move forward. To address this limitation, we propose ReTrail, a training-free decoding method that uses the model's own reflection-token probabilities to guide when and how it revises its reasoning. To determine when to intervene, ReTrail monitors the hesitation mass, defined as the total probability the model assigns to reflection tokens at the end of each reasoning step. This probability measures the model's tendency to reflect, and it remains available even when no reflection token is sampled into the output text. When this signal rises sharply above its earlier values, ReTrail selects an earlier step with high hesitation mass as the starting point for revision. It rolls back to that step and regenerates the continuation. During regeneration, ReTrail first raises reflection-token logits to encourage the model to reconsider that step, and then lowers them to prevent repetitive reflection. The method operates on a single generation trajectory without requiring an external verifier model or multi-candidate search. Across four reasoning models and four benchmarks in mathematics and code generation, ReTrail achieves the highest or tied-highest mean accuracy in all sixteen evaluated settings under identical retained-output token limits. It improves accuracy over unmodified decoding by an average of 3.25 and up to 8.08 percentage points, with a 15.2% average increase in total generated tokens, including tokens discarded during rollback. Component ablations show that rollback improves accuracy while reducing generation length compared to applying reflection bias in place. Furthermore, controlled comparisons of revision timing and restart location show higher mean accuracy with hesitation mass than with representative alternatives like entropy. Code is available at https://anonymous.4open.science/r/retrail-8FC5/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.