Scaling Proactive Compute: When the Agent Anticipates the Task
Abstract
Language model agents benefit from additional test-time compute, yet often remain idle before queries arrive and while tools run. We study how scaling *proactive compute* during these periods improves performance, with three main findings. First, controlled sweeps of preparation and downstream reasoning budgets across nine offline benchmarks show that preparation improves the accuracy–test-time-compute frontier. Reusing studies across questions can also reduce total generated-token costs. In three machine-learning engineering tasks, agents that reason while tools run reach better scores earlier under equal wall-clock budgets. Second, our wiki-style and multi-agent study strategies improve efficiency on Harvey LAB's 134M-token document collection and extend preparation beyond a single context window. Third, our analyses show how preparation reduces downstream reasoning and tool use. On SWE-QA-Pro, a 16k-token study more than halves subsequent reasoning and cuts tool calls roughly tenfold at the same accuracy, even when the solver chooses its own budget. Together, these findings establish proactive compute as an effective complement to test-time scaling and show how agents can turn otherwise idle time into useful preparation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.