Evidence, Not Independence: Reviewer Identity Alone Does Not Verify a Coding Agent's Claims
Abstract
When a coding agent reports a task complete, something else must decide whether to believe it. We test whether reviewer independence, by itself, is sufficient to catch a false completion claim. In a controlled setting where a coding agent repeatedly produces a fix that passes a ticket's stated example but misses an unstated requirement, we vary who reviews (the same model; the same model reframed as reviewing a colleague; a genuinely different model), what they are told to look for, and what evidence they are shown. Reviewer identity alone never catches the omission under a neutral instruction; a stronger instruction narrows the gap for one reviewer pairing but not the other, and even then leaves a case undetected. Showing the reviewer the actual execution output resolves it in every case tested, including when the reviewer is the same agent that wrote the fix, and a visibility-matched check is consistent with the advantage coming from evidence visibility rather than reviewer architecture. A preregistered, fully-crossed replication confirms the pattern at larger scale and surfaces a cost the earlier experiments could not show: the instruction that most improves detection without evidence also raises false alarms on genuinely correct work, an effect evidence only partly offsets. A further probe finds partial evidence is not simply better than none: at an intermediate level of detail, reliability can fall below both a bare pass/fail signal and full output. An unselected random sample of real-repository tasks confirms the underlying failure is not an artifact of how examples were chosen for study. We report this as a mechanism paper, not a general solution: it establishes when execution evidence resolves the failure, leaving how an autonomous system obtains a sufficiently discriminating check without benchmark-held-out tests as open work. We translate the mechanism into a runtime control plane in which completion claims remain proposals until evidence-gated checks admit them into system state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.