When Aggregation Hurts: RESOLVE for Elastic Test-Time Reasoning
Abstract
Scaling test-time compute via structured search—such as Tree-of-Thoughts, Graph-of-Thoughts (GoT), and ensembling—is widely presumed to enhance large language model (LLM) reasoning. In this work, we reveal that expanding search graphs unconditionally often backfires due to Aggregator Poisoning: when candidate traces contain invalid syntax or corrupted logic, presenting them to an LLM synthesizer degrades output quality below single-pass Chain-of-Thought (e.g., on HumanEval, standard GoT collapses from 70.4% to 64.9%). Concurrently, uniform budget expansion wastes compute on straightforward queries, while rigid context-isolation schemes fragment cross-passage dependencies on relational tasks. To resolve these vulnerabilities, we propose RESOLVE (REconciled Synthesis with Optimal-stopping and LLM-gated Verified Exploration). Rather than blind uniform aggregation, RESOLVE maintains full shared context across candidate branches while coordinating: (1)guidance-based width pruning, which syntactically gates out corrupted candidates to shield downstream aggregators from poisoned context, (2)guidance-based depth pruning, dynamically sizing sampling depth via consensus agreement on reasoning or execution feedback on code, and (3)contrastive meta-aggregation, which cross-examines divergent rationales to reconcile conflicting hypotheses. Across 24 benchmark-backbone configurations (, 6 benchmarks, 4 model families), RESOLVE achieves 57.8% overall accuracy compared to 52.1% for GoT, 50.9% for Self-Consistency, 50.4% for ToT, and 50.2% for CoT (outperforming GoT on 20 configurations, tying on 1, and trailing on 3), while reducing token consumption by 40.5% relative to GoT. Under zero-label consensus stopping on natural language and math (4 benchmarks), RESOLVE improves accuracy from 53.4% (GoT) to 56.2% while slashing token costs by 41.1%; on code execution search (2 benchmarks), width gating eliminates aggregator poisoning, elevating accuracy from 49.4% (GoT) to 61.0%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.