Coding Agents Can Deliver What You Check, Not What You Requested
Abstract
Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this evaluation approach has accumulated construction validity problems, and a passing score may not show whether the requested artifact was delivered. In a controlled code-as-spec setup, two production Copilot CLI agents (claude-opus-4.7, gpt-5.5) rebuild a React Fluent UI data table in Angular as a reusable library under a hidden 222-test Playwright oracle across 18 runs and three oracle availability conditions. Alongside the score, we run a mechanical library audit and check each verdict with a no-op ablation. Without the oracle, the library is present but unfinished, revealed by scores. With in-loop verification, the score becomes nearly perfect, but from a demo holding the tested behavior directly, the library left dead or absent. We call this building to the test; the broader self-verification disposition behind both we call validation self-awareness. The agent does not, on its own, validate what it ships as a user would. Prevalence remains an open question across other agents, signals, and model families. Beyond benchmark scores, dispositions like validation self-awareness merit research attention alongside specification gaming and test suite overfitting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.