Beyond Final Confidence: Leveraging Confidence Dynamics for Test-Time Reasoning
Abstract
Test-time scaling improves language model reasoning by sampling and aggregating multiple reasoning trajectories. Recent methods further leverage model confidence to prioritize reliable trajectories. However, final confidence alone can be misleading, as failed trajectories may remain highly confident while repeatedly following an incorrect reasoning direction with little effective self-correction. We refer to this failure mode as high-confidence reasoning stagnation. To distinguish trajectories within the high-confidence regime, we measure the change in token confidence from the penultimate to the final layer and aggregate it over the tail of each trajectory. Among high-confidence samples, stronger inter-layer confidence dynamics are associated with higher accuracy, shorter reasoning trajectories, and less observed stagnation, providing a complementary signal beyond final confidence. Based on this observation, we introduce a training-free aggregation method that jointly filters trajectories using tail confidence and inter-layer confidence dynamics, and then weights the retained trajectories by their confidence dynamics. Across Qwen3-1.7B, Qwen3-8B, and Gemma4-12B, our method consistently improves overall accuracy over standard majority voting and confidence-based aggregation on the evaluated reasoning benchmarks. Further analyses show that these gains remain robust under increasing sampling budgets, while confidence dynamics are most informative when measured across layers close to the final readout.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.