WHERE DOES THE FAULT LIE? TRACEABLE REPAIR FOR CODING AGENTS BEYOND PASS/FAIL
Abstract
Passing tests does not reveal whether a coding agent found the fault, kept edits focused, or avoided repair loops. Across eight models on three external CLI benchmarks, native outcome scores correlate weakly with process proxies from trajectory logs. We introduce TracebenchCLI, a benchmark of 126 CLI repair tasks derived from Codeforces problems. Staged tests and known fault spans support direct measurement of fault attribution, edit locality, progress, wasted turns, and diagnosis. We formalize what the terminal label leaves unresolved through a repair state comprising stage, fault, and accumulated unnecessary work. An information decomposition motivates intervention through action selection. Our controller, **ARC** (Accountable Repair Control), reweights a frozen model's candidates using visible trajectory evidence and requires no training. With fixed settings across five base models, **ARC** raises the Qwen3-32B solve rate from to , reduces wasted turns by up to , and increases explicit diagnosis by up to . Under matched sampling budgets on Qwen2.5-Coder-14B, it solves 10 of 126 tasks against 3 for the base. On three external CLI benchmarks without oracle repair spans, the same controller improves native outcomes and process proxies. In an independent evaluation of 30 pairs matched on terminal outcome, non-author raters prefer **ARC** on 19 pairs and the base on 4, with 7 ties. Code: https://anonymous.4open.science/r/Tracebench-cli/https://anonymous.4open.science/r/Tracebench-cli/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.