acceptodds
Under review as a conference paper at ICLR 2027

FakeWorldEval400: A Contract-Grounded Benchmark for Evaluating Multi-User Coordination in Language Agents

Abstract

Language agents increasingly operate in workflows involving multiple users, ses- sions, and evolving permissions, making reliable coordination essential. However, task completion and response quality alone cannot establish whether an agent addresses the correct recipient, respects disclosure boundaries, and follows cur- rent authority and evidence. To address this evaluation challenge, we introduce FakeWorldEval400, a contract-grounded benchmark for evaluating multi-user co- ordination in language agents. The benchmark comprises 400 manually authored tasks across 22 workflow families and formalizes coordination requirements along five axes: target, scope, authority, source, and time. We develop an architecture- agnostic action interface and a hybrid evaluation protocol that combines deter- ministic contract checks with criterion-level semantic judgments, enabling strict success assessment and trace-based failure localization. Experiments with four mixed-access baselines on 1,600 frozen actor traces yield Hybrid strict pass rates of 37.0%–53.2%. Further analysis of the 207 check-required tasks shows that 81.7%– 89.9% of Rule failures occur while the authored check criterion passes, which localizes most failures to other hard obligations. These results reveal substantial gaps in contract satisfaction within the evaluated multi-user workflows. FakeWorldEval400 provides a reproducible testbed for comparing coordination systems and diagnosing their observable failures under explicit, evolving contracts. Resources are available at https://anonymous.4open.science/r/submission-code-57FE.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.