EAB: An Embodied Agent Benchmark with Anchored Criteria and Paired Interventions
Abstract
Embodied agents fail mid-task: a blocked path, an injected fault, a user interrupt. Outcome-level scoring over dialogue transcripts neither sees these failures nor attributes them. We present EAB, a benchmark for embodied agents at the text-andtool control layer, whose judgment is anchored to execution. Its 400 cases, set in a tour-guide robot domain with 29 skills on 4 MCP servers, 10 real-domain error codes, and 6 mid-task event classes, form 200 strictly matched pairs, each built around one ablated variable, so the within-pair difference estimates that variable’s effect within the benchmark. Verdicts are three-state over 25 criteria aggregated along two axes, four report families and 14 capability units. NA, whether designed away by the case’s criteria scope or never triggered by the episode, is reported as its own verdict and excluded from every denominator. Grading reads execution facts within per-case anchor windows, and the pipeline is validated by golden samples, mutation-tested gates, and a replay ceiling. Evaluating frontier LLM agents under three drive modes (simulator, solo, prescribed), we find that interaction with a user, rather than a skill-level plan, accounts for most of the difficulty: for the three agents run in every mode, only 4.2–8.2% of cases pass all three repeats under interaction, against 38.8–54.0% without a user (0.5–9.8% across all eight agents under interaction). The benchmark then localizes where they break. Terminal success under interaction is comparatively high (43.0% Avg@3 for the reference agent), yet the same episodes satisfy every mounted scored criterion only 8.5% of the time, so the failures are mid-task ones that outcome-scored benchmarks do not see.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.