acceptodds
Under review as a conference paper at ICLR 2027

When Does Self-Evaluation Help in Reasoning? An Anatomy of ToT on Mathematical Tasks

Abstract

Step-wise reasoning, as a fundamental capability of large language models (LLMs), has significantly advanced complex reasoning through decomposing a problem into a sequence of coherent intermediate thoughts conditioned on context information and prior knowledge. As two representatives, sampling-based reasoning paradigms (e.g., Self-Consistency Chain-of-Thought, SC-CoT) achieve impressive performance on tasks with deterministic answers by aggregating multiple trajectories via majority voting, while search-based reasoning paradigms (e.g., Tree-of-Thoughts, ToT) excel on search-intensive tasks that require explicit step-wise evaluations. However, despite explicit step-wise evaluations, ToT often falls short in mathematical reasoning tasks. Rather than simply attributing this problem to unreliable self-evaluations, in this paper, we revisit ToT and explore the properties of its evaluation behaviors from both theoretical and empirical perspectives. Different from SC-CoT, which marginalizes out thoughts and benefits from a concentrated distribution around the correct answer (i.e., low answer diversity), the performance of ToT is correlated with the coverage of branches in a given budget and the distributions of self-evaluations. Motivated by these findings, we propose a tool, entropy-guided adaptive reasoning tree (EGART), to further investigate whether evaluations can provide effective signals to select the promising branches. Overall, our findings uncover the behaviors of ToT in reasoning tasks and provide an insight into why explicit self-evaluation steps fail in search-based stepwise reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.