Rethinking Verification Design in Agent Harnesses
Abstract
Agent harnesses enable large language models to act through tools, observe the results, and iteratively work toward a task goal. Verification checks and revises agents' work to improve output quality. Yet there is limited guidance on how harnesses should organize this effort. We study verification as an axis of test-time scaling by examining when to intervene, who checks the work, and how checking leads to revision. Our framework spans four layers of the agent lifecycle: initial instructions, mid-turn reminders, completion gates, and independent review. We compare 17 designs across seven frontier coding models on Terminal-Bench (TB) and SWE-bench Verified (SWE), and introduce VisualBench for expert evaluation of visual and interactive artifacts. Initial instructions and execution-time reminders yield modest, model-dependent average gains on the coding benchmarks, without substantially increasing average output token use. Completion gates can yield larger gains but may substantially increase cost. Independent review by a critic offers a more favorable accuracy–cost tradeoff, improving average accuracy by – percentage points on TB and – on SWE at – and – baseline cost, respectively. On VisualBench, independent review achieves net preferences of – points over the baseline across its four designs, exceeding the evaluated L0–L2 designs. Additional review rounds or independent critics further improve average TB accuracy by up to percentage points. Together, these results provide empirical guidance for verification design in agent harnesses and support independent review as a practical axis of test-time scaling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.