acceptodds
Under review as a conference paper at ICLR 2027

EA-Bench: Measuring information and action efficiency under imperfect instructions

Abstract

In real-world settings, users often provide incomplete, distracting, or inconsistent information. A useful agent must recognize these imperfections and resolve task-relevant uncertainty through efficient, targeted questioning. Measuring this ability requires separating execution difficulty from the quality of the instruction. To make this distinction measurable, we introduce the Efficient Agent Benchmark (EA-Bench). It holds task records and authoritative policies fixed while systematically varying missing information, redundancy, and conflicting instructions. EA-Bench comprises 270 paired conditions across 18 manually selected synthetic task families. We use task dependency graphs to enforce multi-step execution and counterfactual policy worlds to identify necessary clarifications, creating challenging imperfect instructions consistent with user intent. Experiments on nine mainstream open- and closed-weight model services show that combining multiple imperfections reduces pooled task accuracy from 76.5% to 60.2%. Across missing and mixed conditions, the hardest cases achieve accuracies of only 5.6% and 22.2%. More seriously, even when agents still complete tasks correctly under imperfect information, they make 23.6% more tool calls, request 847.5% more information items, and take 41.0% longer to complete the tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.