On the Limits of Test-Time Compute for Self-Evaluation
Abstract
Self-verification and aggregation have emerged as popular approaches for bridging the gap between and . However, increasing the number of candidates does not consistently narrow the gap between final-answer accuracy and . To understand this limitation, we use ranking error to measure how reliably models distinguish their correct generations from incorrect ones. We show that ranking error admits a bias-variance decomposition, with both components shaped by algorithmic design choices and inference compute. Across our experiments, additional compute generally reduces variance, while bias initially decreases before approaching a plateau. Guided by this analysis, we introduce Scorpio, a simple and general method that jointly scores randomly grouped candidates. Across math, code, and puzzle benchmarks, Scorpio improves verification, tree search, and aggregation. For verification, Scorpio improves accuracy over the strongest baseline by up to percentage points while matching or exceeding that baseline’s peak accuracy with – lower total inference cost. For tree search, using Scorpio to score partial solutions and select the final answer improves accuracy by up to 10 percentage points over prior methods. Finally, Scorpio improves the scaling of aggregation-based methods by percentage points on coding tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.