acceptodds
Under review as a conference paper at ICLR 2027

Certified Selective Automation of LLM Agent Evaluation

Abstract

Evaluating LLM agents still ends with a human reading trajectories, because the automatic judges that could replace this reading come with no guarantee on how often they are wrong. We ask the operational question directly: what fraction of agent-trajectory evaluation can be handed to a judge, with a certificate that the error rate among auto-decided trajectories stays below a budget ? Answering it on agent corpora runs into a structural obstacle: one task is attempted by many agents, so trajectories arrive in correlated clusters, and the i.i.d. certificates used by existing selective-judging methods can overstate what is safe: on one of our corpora a naive certificate claims 98% automation while its realized error exceeds the budget in 17.5% of task resamples. We introduce a task-level bootstrap certificate that is valid in every regime we test, synthetic and real, while matching the naive certificate's coverage, and we show that the finite-sample cluster-valid alternatives certify nothing at realistic task counts. With the certificate fixed, we study the judge: a 4B logprob judge, fine-tuned and then trained with reject-weighted GRPO, certifies 0.30-0.59 of evaluation on tool-use and web corpora at , is the only judge (among strongly elicited frontier models) that certifies on both headline corpora, and costs two orders of magnitude less per decision than the one frontier configuration that beats it. Certified coverage turns out to be predictable before training from base-rate and discrimination alone (leave-one-corpus-out ), which tells practitioners when certification is achievable and when reinforcement learning will help. Finally, the certificate does double duty as a self-training filter: harvesting pseudo-labels only inside certified regions bounds their contamination by by construction (realized: 0.000-0.041 across six harvests), lets a judge enter an unseen domain with zero target-domain training labels at in-domain strength, and improves the strongest corpus beyond its best supervised judge.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.