To Aggregate or Not to Aggregate? Test-Time Aggregation Beyond Verifier-Friendly Benchmarks
Abstract
Aggregation-based test-time scaling has produced strong gains on competition mathematics and code generation, but it remains unclear whether those gains transfer beyond verifier-friendly benchmarks, under which task conditions aggregation helps, and when one-step aggregation is sufficient relative to recursive aggregation. We study these questions across structured reasoning, knowledge-intensive reasoning, and medical reasoning, spanning proof-style mathematics, expert-level STEM reasoning, social-science knowledge tasks, BrowseComp-style information seeking, and both tool-free and tool-integrated regimes. Aggregation is effective when sampled trajectories contain recoverably complementary information: complementary reasoning progress in structured reasoning, or complementary retrieved evidence in tool-integrated knowledge/evidence-seeking. Structured reasoning and tool-integrated knowledge/evidence-seeking recover 48% and 57% of available headroom on average, whereas medical reasoning without tools recovers only 21% despite a comparable diversity window. Tool use improves medical base performance, but aggregation remains weak because trajectories more often reflect competing clinical interpretations than composable intermediate progress. Within favorable regimes, aggregation type also matters: RSA is most useful for open-ended proof generation, whereas SSA captures most of the gain in tool-integrated knowledge/evidence-seeking at lower cost. Aggregation value therefore depends jointly on task structure, tool access, and aggregation type rather than following a uniform test-time scaling law across domains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.