acceptodds
Under review as a conference paper at ICLR 2027

Agentifying AI Evaluation: What to Evaluate, When to Stop, and How to Estimate

Abstract

Evaluating modern AI agents is increasingly costly and time-consuming, as benchmarking tasks can require long sequences of model calls and tool interactions. Can we reduce evaluation time and token cost without sacrificing model evaluation accuracy? Recent methods save cost either by selecting a subset of tasks to run or by stopping trajectories early based on partial execution histories, but generally trade lower cost for reduced evaluation accuracy. We introduce an agentic evaluation framework that jointly and adaptively learns what to evaluate and when to stop, with a tailored estimator that preserves evaluation accuracy. The system jointly learns what tasks to evaluate and when to stop execution, while probabilistically randomizing and adapting these decisions to enable statistical correction. By "learning what tasks to evaluate," it uses historical evaluation traces to identify which benchmark tasks are most informative for estimating agent performance. By "deciding when to stop," it actively monitors each trajectory and probabilistically continues only when further execution is expected to provide sufficient additional information about the eventual outcome. We pair this agentic system with a tailored estimator that corrects for adaptive task selection and stopping. Our theoretical analysis shows that the resulting framework preserves model evaluation accuracy while improving precision for a given evaluation budget. Experiments on benchmarking data show that our approach reduces both evaluation time and token cost while preserving accurate model evaluation results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.