acceptodds
Under review as a conference paper at ICLR 2027

Built on Broken Assumptions (BOBA): Why Coding Agents Cannot Be Left Alone

Abstract

Coding agents are autonomous developers. As such, they can help not only when each change is reviewed before the next is requested, but also when a whole sequence of requests is left to them, each built on the last. Although agents rarely notice their own mistakes and report wrong work as complete, evaluation has focused on a single request from human-written code, verified once. In this work, we build BOBA-BENCH, 280 chains of 20 consecutive requests from 140 real repositories, and compare agents on single requests and across chains. All five frontier agents we test complete far fewer chains than requests: 48% on average against 94–98%. Two matched comparisons show that frontier agents’ verified code is as good a base as human code, and that the degradation comes from building on unverified code. Agents submit work as complete without confirming it and build on the broken state unaware, whereas one external “that’s wrong” and a repair round lift chain completion by 8 to 16 points. In simpler terms, what limits coding agents left alone is not the code they write but the mistakes they cannot see: without verification, every request that follows is built on broken assumptions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.