Stop Stealing, Stop Attacking: Countering Benchmark Hacking with Kernel-Enforced Contracts Outside the Agent's Reach
Abstract
Instead of solving a benchmark task, an agent can steal the answer—and attack real systems to obtain answers. OpenAI reported evaluated agents compromising Hugging Face's production infrastructure while seeking test solutions. Benchmark hacking thus threatens both the validity of evaluation scores and the security of systems outside the benchmark. Existing controls cannot distinguish protected answers from legitimate references on the same host, or unauthorized operations from permitted ones. Post hoc audits cannot undo exposure or external effects. We give each evaluation run a bidirectional contract listing which answers may not come in and which operations may not go out. Written in a declarative domain-specific language (**DSL**), the contract also states the evidence a score requires and is frozen before each run. To better prevent the agent from escaping, we choose to have the host operating-system kernel, which is outside the agent's reach, assist with the checks. We reduced successful answer theft by **81.8%** on 60 SWE-bench Verified and Pro tasks, and the agents' escape rate by **91.3%** in our attack study.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.