How Does Test-Time Compute Affect Agentic Forecasting?
Abstract
Test-time compute scaling has been explored extensively in easily verifiable domains such as math and code. However, much less is known about how the test-time budget should be spent for forecasting future events, where no test-time verifier is available. We study the problem of compute allocation for *agentic forecasting* by comparing three multi-agent policies that spend inference-time compute in different ways: static depth (*Predictor-Critic*), static breadth (*Ensemble + Aggregator*), and adaptive routing (*Hierarchical Orchestrator*). We perform a contamination-controlled backtesting evaluation on label-balanced binary questions from ForecastBench partitions, all resolved strictly after every base model's knowledge cutoff, under a shared web-filtering layer that screens retrieved pages for forecast leakage and post-resolution contamination. In the gpt-5.4-mini sweep, adaptive routing occupies the entire cost-accuracy Pareto frontier: the top Orchestrator reaches accuracy at USD per question, while the best Ensemble and Predictor-Critic reach at USD per question and at USD per question, respectively. The same Pareto ordering replicates on the gpt-5.4-nano sweep, where the Orchestrator is about cheaper than the top Ensemble, with no statistically significant accuracy gap. The Orchestrator's cost-quality advantage further extends to other base models, including Claude Sonnet 4.5, DeepSeek v4 Flash, and Gemini 2.5 Flash, so the result is not a single-model artifact. A two-stage mechanism diagnostic explains this win as *selective delegation*: on three of five base models, the Orchestrator allocates more compute to questions where its cheap direct baseline is uncertain, and this uncertainty can predict where extra compute helps. Adaptive routing wins not by spending less on every question, but by concentrating compute on the ones where it pays off. We release the benchmark, architecture implementations, pipeline, and analysis code.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.