acceptodds
Under review as a conference paper at ICLR 2027

Who Catches What? A Trace-Based Study of Heterogeneous LLM Review and Repair

Abstract

A common safeguard for software-engineering agents is to have a second model review the first model’s work. Existing measurements of this practice, however, come largely from benchmarks and simulated red-teaming. We study it in a production deployment using ten weeks of event logs from KISS Sorcar, where an Anthropic model implemented changes in a working repository and an OpenAI model reviewed them read-only under a budget cap. The corpus contains 486 tasks and 1.1M events, including tool calls, model switches, and per-step costs. The logs show that reviewer access rarely resulted in writes: among 2,164 reviewer-attributable calls, only 3 (0.14%) performed a write, and none modified tracked source files. On the 42 tasks for which review spending can be isolated, review accounted for 9.3% of total task cost. Task narratives contain 1,943 reported reviewer findings and 1,779 corresponding fixes, but these counts are self-reported by the builder; an independent re-coding finds roughly one-fifth fewer findings. In a 60-task re-coded sample, both readers identify 5 tasks where review exposed a problem that the builder’s tests had not caught, while the second reader identifies 7. Three case studies from the full corpus illustrate such findings in detail, including a code-execution path from model output and a repair that introduced a linearizability race. These observations characterize how model-based review is used in production and what it costs, but they do not establish its causal benefit because the deployment has no comparison arm. We release the trace corpus and analysis scripts to support further study.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.