acceptodds
Under review as a conference paper at ICLR 2027

MACE-Bench: Measuring Available and Realized Parallelism in Multi-Agent Software Engineering

Abstract

Current multi-agent evaluations report whether a task was solved, not whether it offered parallel work that a multi-agent system (MAS) could have used. We introduce MACE-Bench, which makes that work executable and observable. Each of its 104 long-horizon software tasks carries an executable reference DAG: one feasible decomposition of the work, recovered from a successful trajectory and tested by replay under the task’s original verifier. Frontier model–harness systems run each task as a single agent, as the harness’s own subagents, as those subagents given the reference plan, and as a deterministic scheduler that launches one agent per ready unit. When, without a plan, the model decides whether to delegate, cost rises up to 2.06× with no detected speedup on any system. When the runtime launches ready units, tasks solved by both conditions finish 1.77× to 2.02× faster at 1.00× to 1.32× the single agent’s cost, without solving fewer tasks; the scheduler’s plan is solution-derived and the scheduler verifies less, so part of this gain is not parallelism. Trace analysis finds that launching workers is cheap and locates the lost time in parent reasoning before delegation, retained critical work, and serial verification after the workers finish.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.