acceptodds
Under review as a conference paper at ICLR 2027

FWCollab: Benchmarking Structured Multi-Agent Collaboration with Executable Dependency Graphs

Abstract

Terminal success does not describe how a team establishes, maintains, and recovers the intermediate conditions required by its task. We introduce FWCollab, a symbolic two-agent benchmark that operationalizes these cross-role dependencies as private, benchmark-authored executable graphs bound to authoritative environment state. A deterministic, replay-authenticated evaluator measures progress, handoffs, condition regressions, and structural blockers without a language-model judge. The construction library contains 72 Core tasks; evaluation uses a selected 24-task Core subset plus information, bounded-synchronization, and parallel-join tracks, totaling 52 tasks and 936 episodes across six models. Under the tested budgets, intermediate dependency satisfaction and environment completion can differ substantially; these differences characterize execution rather than establish general collaboration quality. A separate paired communication study is analyzed within its own fixed-budget protocol, with interface diagnostics and sensitivity to how controller releases are counted. Conformance and alternative-execution checks validate implementation against the declared specification, not human construct validity. A repair audit identifies historical credit for already-completed stages; requiring new progress removes clear evidence of either selector’s superiority. FWCollab provides a reproducible instrument for examining specified cross-role dependencies, with explicit boundaries between executable measurement, downstream utility, and capability comparison.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.