acceptodds
Under review as a conference paper at ICLR 2027

The Grader Is an Intervention: Evaluator Non-Interference for Stateful Agents

Abstract

Benchmarks for AI agents score an agent by checking the environment it leaves behind, assuming the check changes nothing. Training, search and early stopping now grade episodes in progress or more than once. We show that a grader can score the current state correctly and still change a later score. We call the requirement that extra grading leave recorded outcomes unchanged evaluator non-interference, and test it in shipped code. Computing τ-bench’s reward writes the correct answer into the environment, so a public training wrapper that grades twice reports 127 passes where the benchmark reports 22. On AppWorld, grading after every API call turns 17 of 142 published successes into failures, and the benchmark’s own analysis script destroys 8. There, one grading call overwrites shared state in two ways, and undoing each repairs exactly its own cases, even on held-out tasks. With WebArena’s original evaluator, though not its PyPI port, Browser-Gym’s adapter lets the grader move the agent’s page. In a disposable copy of the environment’s process, grading changes no replayed outcome beyond run-to-run noise. Grading is therefore part of the experiment, and benchmarks should state when their grader runs and what reads the state it writes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.