acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Stability for Reasoning Paths in Large Language Models

Abstract

Large language models (LLMs) demonstrate remarkable capabilities in complex reasoning tasks, where test-time compute has emerged as a crucial means for performance enhancement. Existing methods mainly favor outputs supported by answer consistency or high token confidence, implicitly treating stable reasoning paths as more reliable. However, these approaches provide limited guidance on when stability is a reliable selection signal. In this study, we analyze how the relationship between token stability and correctness varies with the model's confidence in its reasoning path. Our analysis shows that stability is informative when the model is confident, whereas unstable paths should not be uniformly discarded when the model is uncertain. Motivated by this finding, we introduce the (), which combines macro-level verbalized confidence with micro-level token confidence fluctuations to filter reasoning paths across confidence regimes. Comprehensive evaluations across eight diverse benchmarks spanning mathematical, multidisciplinary, code, and multimodal reasoning tasks on three model families (Qwen, Llama, and InternLM) demonstrate that VeTo achieves the highest average accuracy within each model family among all compared strategies. Crucially, our method scales robustly with inference budgets from to and across model sizes from 4B to 14B, while also yielding competitive calibration performance in mathematical reasoning. Overall, our findings provide a new perspective on reasoning path selection by interpreting stability and instability according to model confidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.