Not All Thoughts Are Worth Thinking: Reliability-Aware Test-Time Scaling
Abstract
Test-time scaling improves LLM reasoning, but concept-guided inference can waste compute on weak or redundant solving directions. We study pre-expansion concept allocation: which explored concepts should receive solution-generation compute? We introduce RATTS, a Reliability-Aware Test-Time Scaling method that combines calibrated predictions of judged concept usefulness with semantic soft coverage. Its objective favors useful, complementary concepts, is monotone submodular, and admits a greedy approximation guarantee; under an explicit concept-success model, we further derive a conditional bound connecting the surrogate to candidate recall. Across Math500, GPQA Diamond, and HumanEval+ with three generators, RATTS achieves macro-average pass-any at and , versus for score-only selection and for uniform-weight semantic coverage. RATTS retains of full-expansion pass-any while reducing exploration-plus-expansion tokens by . Macro-average final accuracy is versus for full expansion, with a paired difference of pp ( CI: ). To enable reproduction of these results, we will release the code and experimental artifacts upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.