Who Catches What? A Trace-Based Study of Heterogeneous LLM Review and Repair
Abstract
A common safeguard for software-engineering agents is to have a second model review the first model’s work. Existing measurements of this practice, however, come largely from benchmarks and simulated red-teaming. We study it in a production deployment using ten weeks of event logs from KISS Sorcar, where an Anthropic model implemented changes in a working repository and an OpenAI model reviewed them read-only under a budget cap. The corpus contains 486 tasks and 1.1M events, including tool calls, model switches, and per-step costs. The logs show that reviewer access rarely resulted in writes: among 2,164 reviewer-attributable calls, only 3 (0.14%) performed a write, and none modified tracked source files. On the 42 tasks for which review spending can be isolated, review accounted for 9.3% of total task cost. Task narratives contain 1,943 reported reviewer findings and 1,779 corresponding fixes, but these counts are self-reported by the builder; an independent re-coding finds roughly one-fifth fewer findings. In a 60-task re-coded sample, both readers identify 5 tasks where review exposed a problem that the builder’s tests had not caught, while the second reader identifies 7. Three case studies from the full corpus illustrate such findings in detail, including a code-execution path from model output and a repair that introduced a linearizability race. These observations characterize how model-based review is used in production and what it costs, but they do not establish its causal benefit because the deployment has no comparison arm. We release the trace corpus and analysis scripts to support further study.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.