Rethinking Stability for Reasoning Paths in Large Language Models
Abstract
Large language models (LLMs) demonstrate remarkable capabilities in complex reasoning tasks, where test-time compute has emerged as a crucial means for performance enhancement. Existing methods mainly favor outputs supported by answer consistency or high token confidence, implicitly treating stable reasoning paths as more reliable. However, these approaches provide limited guidance on when stability is a reliable selection signal. In this study, we analyze how the relationship between token stability and correctness varies with the model's confidence in its reasoning path. Our analysis shows that stability is informative when the model is confident, whereas unstable paths should not be uniformly discarded when the model is uncertain. Motivated by this finding, we introduce the (), which combines macro-level verbalized confidence with micro-level token confidence fluctuations to filter reasoning paths across confidence regimes. Comprehensive evaluations across eight diverse benchmarks spanning mathematical, multidisciplinary, code, and multimodal reasoning tasks on three model families (Qwen, Llama, and InternLM) demonstrate that VeTo achieves the highest average accuracy within each model family among all compared strategies. Crucially, our method scales robustly with inference budgets from to and across model sizes from 4B to 14B, while also yielding competitive calibration performance in mathematical reasoning. Overall, our findings provide a new perspective on reasoning path selection by interpreting stability and instability according to model confidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.