Benchmarking LLM Agents on Pushing the Test-Time Reward-Cost Frontier
Abstract
LLM agents can now autonomously form hypotheses, modify code, run experiments, and interpret results—a workflow commonly known as autoresearch. A prominent goal of this paradigm is to use LLM agents to improve LLMs themselves. Existing autoresearch testbeds typically pursue this goal through model training. This makes each research iteration slow and expensive, leaving agents with limited opportunities to develop methods that are practical and transferable to larger-scale settings within a fixed time budget. Moreover, because agents are often given broad control over data, training pipelines, and engineering choices, final performance may depend as much on auxiliary capabilities—such as finding benchmark-relevant training data—as on the ability to conduct research. We instead introduce a controlled benchmark for autoresearch at test time. It asks an agent to push the reward–cost frontier of a frozen LLM on agentic tasks solely by reorganising its inference-time computation: achieving a higher reward at the same cost or the same reward at a lower cost. This setting enables faster research iterations, and any method discovered can be applied directly to already-trained models without additional training. We build a controlled evaluation on -bench, Terminal-Bench 2.1, and SWE-bench Verified, using a small development set for autoresearch and a separate held-out set for evaluation. During autoresearch, coding agents may change only how the given LLM is used at inference time. We find that (1) frontier agents cannot yet push the reward–cost frontier of their own backbone models, even when they appear to do so on the development set; (2) different agents converge to similar methods; and (3) agents retain the broad research direction they choose at the outset, even when given access to extensive relevant literature.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.