acceptodds
Under review as a conference paper at ICLR 2027

Gut Check: Benchmarking Coding Agent Observation Judgments via DoubleTakeBench

Abstract

Coding agents solve complex software tasks by iteratively interpreting tool observations, but deliberating over execution logs with slow reasoning models is prohibitively expensive, while unverified errors trigger costly cascading detours averaging 6.2 unguided tool calls. Lightweight System 1 models promise a fast alternative, evaluating execution claims in a single forward pass, yet prior evaluations measure agreement with larger models on synthetic workflows rather than execution ground truth. We introduce DoubleTakeBench, an execution-grounded benchmark of 525 decision-time observation judgments mined from revision cycles across 476 real agent trajectories and 11 model backbones. By isolating local claim verification under bounded decision-time evidence, our benchmark provides a standardized testbed for fast judgment in software environments. Evaluating 14 System 1 systems against eight frontier language models shows that top fast judges reach accuracy competitive with non-reasoning language models (61.5%–64.3%) but exhibit pervasive affirmation bias on false statements and sensitivity to option order. Finally, analysis shows that calibrated confidence gating helps identify reliable predictions, demonstrating how fast monitors can avert unnecessary agent detours.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.