When Passing Checks Are Not Enough: Separating Evidence Reports from Outcome Judgments in Coding Agents
Abstract
Coding agents can pass their own tests and still falsely claim completion. We distinguish reporting verification events accurately from judging whether those events justify confidence in success. We introduce an E/J completion interface and verdict training on agent-generated trajectories. Evidence reports (E) are supervised by execution records, while outcome verdicts (J) predict hidden evaluation outcomes. On SWE-bench Verified, an SFT baseline with this interface performs verification but remains overconfident in its verdicts. We train J on the baseline's trajectories with two targets: Hard-J uses observed outcomes, and Soft-J blends outcomes with a difficulty-conditioned estimate of evidence reliability. The two targets have complementary strengths. On fixed development histories, Hard-J raises AUC from 0.590 to 0.787, while Soft-J reduces expected calibration error from 0.294 to 0.083. On 72 held-out tasks, Hard-J makes the fewest false-positive verdicts after passing checks (28%, versus 39% for Soft-J and 83% for SFT), while Soft-J has lower calibration error and Brier score than independently calibrated Hard-J on every trajectory source, with comparable discrimination among passing-check episodes. Ablations and report interventions show that verdicts condition on reported evidence, and directly incentivizing positive self-reports induces reporting hallucinations, reinforcing the need to supervise process fidelity separately from outcome judgment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.