HackDetect: Benchmark Validity in the Age of Agentic AI
Abstract
Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the benchmark keeps the intended capability necessary for success. Recent reward-hacking benchmarks and system reports show that agents can instead recover public solutions, read evaluation artifacts, infer generator structure, manipulate feedback, or benefit from invalid scoring paths; existing responses do not provide a common procedure for attributing these shortcuts and quantifying their effect across benchmarks. We formalize benchmark validity and introduce HackDetect, a post-hoc audit in which an LLM judge identifies the shortcut a benchmark exposes (Exposure) and whether the agent used it (Engagement); comparing scores with and without the shortcut then determines whether the reported score is misleading (Mislead). Across 1,404 records from 11 public agent benchmarks, agents use benchmark-exposed shortcuts in 67.0% of Frontier Science traces, and 24 of 36 AutoLab traces contain a supported exposure. Five audited cases show score gaps of – between runs with and without the shortcut, demonstrating that benchmark reports should provide evidence that scores reflect the intended capability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.