Gut Check: Benchmarking Coding Agent Observation Judgments via DoubleTakeBench
Abstract
Coding agents solve complex software tasks by iteratively interpreting tool observations, but deliberating over execution logs with slow reasoning models is prohibitively expensive, while unverified errors trigger costly cascading detours averaging 6.2 unguided tool calls. Lightweight System 1 models promise a fast alternative, evaluating execution claims in a single forward pass, yet prior evaluations measure agreement with larger models on synthetic workflows rather than execution ground truth. We introduce DoubleTakeBench, an execution-grounded benchmark of 525 decision-time observation judgments mined from revision cycles across 476 real agent trajectories and 11 model backbones. By isolating local claim verification under bounded decision-time evidence, our benchmark provides a standardized testbed for fast judgment in software environments. Evaluating 14 System 1 systems against eight frontier language models shows that top fast judges reach accuracy competitive with non-reasoning language models (61.5%–64.3%) but exhibit pervasive affirmation bias on false statements and sensitivity to option order. Finally, analysis shows that calibrated confidence gating helps identify reliable predictions, demonstrating how fast monitors can avert unnecessary agent detours.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.