acceptodds
Under review as a conference paper at ICLR 2027

REVIBE: Quantitatively Evaluating Multi-Agent Vibe Coding on Real-World Software Tasks

Abstract

Vibe coding is reshaping software development, and multi-agent teams promise to compress delivery time by dividing labor across parallel workers. Yet existing multi-agent benchmarks rely on simulated settings that entangle reasoning with communication, only reward coarse task completion, and ignore the time and monetary costs that matter most in practice. We introduce REVIBE, a benchmark that quantitatively evaluates multi-agent vibe coding on real-world software tasks. REVIBE crosses 10 full-stack projects, each given to agents only as a natural-language requirement document and graded by a weighted 100-point rubric, with 10 explicitly defined collaboration modes: complete organizational policies that fix how a four-agent team allocates ownership, concurrency, decision authority, and handoff artifacts, while the model, team size, tools, round budget, and evaluator stay fixed. Two components make the grid runnable: LEGOGENT executes each mode through periodic sync communication and a native CI/CD pipeline, and TAGENT discovers and probes each unseen deployment to report a functional score together with latency and prefix-cached token cost. Across 100 project–mode configurations and more than 300 runs, the collaboration mode moves the average best score about as much as the choice among three models does, and far more on a given task: with the project and model fixed, switching the mode shifts the score by up to 40 points and the time to the best round by up to 4.1. No mode dominates: six modes lie within two points of each other across projects, and the winning mode depends on the project's requirement structure. REVIBE makes these speed, cost, and quality trade-offs measurable and reproducible. Code and benchmark are available at https://anonymous.4open.science/r/revibe_repo-8828/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.