Same Uncertainty, Different Remedies: Failure-Aware Routing of Test-Time Compute
Abstract
Adaptive test-time scaling decides where and how much inference compute to spend, but after an initial sampling batch the system must decide what the next computation should do. Scalar uncertainty can signal risk, yet it does not specify whether to retain the majority, verify competing candidates, or abstain. We study whether answer-distribution structure provides a routing state using answer counts alone. Across 1,577 GPQA question–model pairs, the distributions admit operational regimes with different candidate concentration and intervention profiles. This motivates a failure-aware router that keeps majority voting outside a contestation gate, verifies concentrated top-two contests with a sampler-specific screened verifier, and abstains on scattered cases. In the primary Qwen3-32B/GPQA setting, verification gains percentage points on 17 two-mode flags. With 51 verifier calls, full-coverage accuracy rises from 54.5% to 58.7%; abstaining on 11 scattered flags yields 61.4% accuracy at 92.3% coverage. The router has the largest observed equal-budget point estimate, although its advantage over runner-up-share allocation is not statistically established. Verification utility also varies across sampler–verifier pairs. These results suggest that uncertainty identifies risk, structure identifies candidate availability, and realizing that opportunity additionally requires a competent resolver. Routing should therefore follow the observed failure configuration and screen the resolver it depends on.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.