acceptodds
Under review as a conference paper at ICLR 2027

The Evaluator Also Varies: Localizing Nondeterminism in Interactive Agent Benchmarks

Abstract

Interactive agent benchmarks can use large language models both as agents and as components of the evaluation system. When repeated evaluations disagree, the variation may originate in the simulated user or environment before it appears in the agent’s behavior. We introduce a first-divergence audit that identifies the earliest disagreement between replicated interactions with identical preceding recorded histories, locating it in the agent, simulated user, or environment. Across 100 airline and retail tasks on τ²-bench, five greedy vLLM replicates per task reveal transcript divergence on all tasks, but reward variation on only 20. The first discrepancy appears in the simulated user on 83 of the 100 tasks. In banking, disagreements persist in environment outputs after prefix caching is disabled on the agent and simulator. Comparisons of serving configurations reveal different patterns of residual divergence. A complementary ALFWorld experiment shows that preceding serving history can change trajectories without changing aggregate success, while cache resets restore agreement in the tested conditions. Together, these results show that stable benchmark rewards can conceal an unstable evaluation process. Reproducibility must therefore be assessed across the complete interaction, including the simulated user, environment, and serving protocol.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.