Benchmarks as Software: Toward Continuous Benchmarks from Terminal-Bench
Abstract
Coding agents are among the most economically consequential deployments of LLMs to date, making reliable evaluation increasingly important. Yet even benchmarks that are carefully validated at release can become stale. As models improve, they can discover valid solutions that existing verifiers were not designed to recognize, while changes to evaluation infrastructure can introduce new sources of error. We call the resulting accumulation of defects benchmark maintenance debt: a growing gap between benchmark scores and the capabilities they intend to measure. We introduce a Continuous-Validation Workflow for agentic benchmarks that combines community reports, task artifacts, agent traces, automated oracle checks, and human-gated repairs. We apply this workflow to Terminal-Bench 2.0, a benchmark that underwent extensive human and LM-assisted review and red-teaming before release. Our audit identifies issues in roughly 30% of the benchmark, and we release the resulting repairs as Terminal-Bench 2.1. These repairs change evaluation results: across agents, accuracy shifts by as much as 12 percentage points and leaderboard positions move by up to three ranks, even as the overall ranking remains highly correlated (Spearman ). We also show the workflow generalizes to tasks in other agentic benchmarks. Beyond accuracy, we evaluate how resource allocation shapes benchmark conclusions. We show that fixed wall-clock and output-token budgets can materially change apparent agent performance: some large harness gaps at the default timeout shrink substantially when agents are given more time. These results show that benchmark maintenance is not only about repairing task artifacts; resource budgets are part of the measurement specification and must be validated alongside the tasks themselves. More broadly, benchmark maintenance debt is not just accumulated implementation bugs: our results show that as models improve, both evaluation assumptions and resource settings may need to be revalidated.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.