acceptodds
Under review as a conference paper at ICLR 2027

Is This Agent Evaluable? Certifying Verdicts About LLM Agents Under Nondeterminism, Oracle and Coverage Obstructions

Abstract

A benchmark score is an instrument reading. The instrument has four stages: a policy that samples actions, an environment that carries state, an oracle that renders a verdict, and a task suite somebody chose. We define evaluability as the existence of a procedure that settles a claim from the available evidence at a stated error level. We name five ways it fails, and measure each on live agents across three domains with unrelated checker designs. Four results. Swapping the tasks moves a score 6.2 times more than rerunning them, so a three-seed error bar prices the cheaper axis. Whether reruns settle a task is a fact about the pass rates, not the checker, so a fixed suite settles less as agents improve. Judges agree at 0.761 on benchmark obligations but only 0.513 on capabilities that real agents declare. And cheatability belongs to the interface: our own harness leaks every reference answer, making 0.608 of our suite passable with no task competence, none of it reachable by the five browser primitives our agent had. Of nine pre-registered predictions, 3 survive. Where an obstruction binds, the honest output is undecided with the reason named, rather than a score.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.