acceptodds
Under review as a conference paper at ICLR 2027

CommandBench: Toward Diverse, Challenging, and Contamination-Resistant Repository Generation

Abstract

Coding agents are advancing quickly, and automating the entire software engineering process is becoming an increasingly important setting. Existing benchmarks, however, evaluate this setting only weakly: open-source repositories, the single most important source from which benchmarks are built, have been widely absorbed into the training corpora of frontier LLMs, and it is no longer possible to tell whether a high score reflects genuine ability or a replay of the training corpus. Having experts write tasks by hand does resist contamination effectively, but it runs into a clear bottleneck in scale and diversity. We therefore present CommandBench, which asks an agent to implement a complete repository from scratch and encapsulate it as a set of CLI commands. CommandBench is resistant to contamination in three ways: (i) lexical obfuscation, which makes a located repository unusable as it stands; (ii) cross-language implementation, which blocks verbatim copying; and (iii) unmet-specification tests, which target behaviour that the specification requires but that the original repository does not yet achieve. Most distinctively, CommandBench guarantees that an implementation that reproduces the source repository's behaviour scores exactly zero, so that faithfully replaying the source implementation earns no credit by construction. CommandBench is also markedly more diverse and more challenging: its tasks are synthesised from 138 distinct repositories, covering 7 programming languages, with tasks from outside computer science accounting for 48.6% of the benchmark. On CommandBench, models pass 89% of the met-specification tests on average, yet none passes more than 63% of the unmet-specification ones, and the best score is only 58.93 out of 100.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.