TITRATE : Evaluating LLM agents with fewer executions
Abstract
Agentic benchmarks run an agent through multi-step tasks in real environments. Each run costs money and compute, and its outcome is random. Existing efficientevaluation methods fix a budget, typically 10% of the tasks, and report whatever accuracy results. An evaluator usually needs the opposite: given a target accuracy, how many runs are required? With 10% of the tasks, the share of agents scored within four points of their full-benchmark score ranges from 39% to 83% across five benchmarks, so the same budget buys very different reliability. We propose T ITRATE, which derives the number of runs from the tolerance and adapts item response theory to agentic benchmarks in three ways. Task parameters are fitted on the few dozen agents a benchmark has already evaluated, so T ITRATE shrinks the estimated discriminations to keep tasks that look informative by chance from being selected first. An agent–task random effect separates an agent’s lasting advantage on a task from run-to-run randomness, which tells T ITRATE which tasks are worth rerunning. Because the model’s own error estimate is optimistic, T ITRATE scales it by a factor learned by replaying the procedure on those agents, which turns a tolerance into a task count. Asked to place 95% of agents within four points of their full-benchmark score, T ITRATE places 91% to 97% on all five benchmarks under random splits. With 10% of the tasks, it estimates the full-benchmark score at least as accurately as six existing methods on every benchmark, and choosing which tasks to rerun matches uniform rerunning with about 20% fewer runs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.