OpenCollab: A Multi-Agent Coding Framework with Programmable Collaboration and Controllable Runtime
Abstract
Multi-agent coding systems tackle complex software engineering tasks through specialized roles, typically measuring their collaboration benefits against a single-agent baseline. However, existing evaluations blindly assume the configured organization is faithfully executed, whereas agents often solve tasks entirely in isolation without any actual delegation. Furthermore, confounding factors like models, toolsets, and token budgets frequently vary across systems, preventing observed gains from being cleanly attributed to the organization itself. To address this, we introduce OpenCollab, a framework for controlled multi-agent evaluation. OpenCollab defines adherence to quantify whether the declared organization is actually realized during execution. It provides a unified runtime to hold all key experimental factors strictly constant, and records a fine-grained event stream to enable per-run verification between assigned and realized organizations. On SWE-bench Verified, we find that adherence is strictly zero across 40 tasks when the model is free to choose whether to collaborate. By altering only one closing instruction in the role prompt to mandate delegation, adherence jumps to 0.925 on the same tasks. These findings demonstrate that multi-agent performance claims are uninterpretable without measuring adherence, establishing OpenCollab as a necessary infrastructure for rigorous, causal multi-agent comparisons.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.