acceptodds
Under review as a conference paper at ICLR 2027

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

Abstract

As autonomous coding agents produce more code than any developer can review, oversight collapses onto a single surface: the automated test suite. Reward hacking naturally arises in this setup, as the agent optimizes over the target of passing more automated test, the test suite cease to be a good measure for the developer to review the real progress of the agents. We introduce SpecBench, a benchmark of 30 systems-level programming tasks, spanning reference implementations from a 1,500-line JSON parser to a 110,000-line operating-system kernel. Each task includes two test suites: (i) visible validation tests that exercise specified features in isolation, and (ii) held-out tests that compose those same features to simulate real-world usage. Crucially, these held-out suites introduce no requirements beyond the visible validation suites, a genuinely implementation will inherently pass the validation suites as well as the held-out suites. Thus we can use the gap in pass rates on these two suites to quantify reward hacking. Large scale experiments reveal a consistent pattern: every frontier agent saturates the visible suite. Yet reward hacking persists, with smaller models demonstrating larger gaps between test suites. The gap also scales sharply with task length: the gap grows by 28 percentage points for every tenfold increase in code size. Failures range from subtle feature isolation to deliberate exploits, including a 2,900-line hash-table "compiler" that memorizes test inputs. SpecBench offers a principled testbed for measuring whether coding agents build genuine working systems or merely game the test suites developers hand them.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.