ReqAuditBench: Benchmarking Coding Agents on Proactive Discovery of Requirement–Code Inconsistency
Abstract
Coding agents increasingly write code for software projects, but checking whether the code is consistent with documentation is left to human reviewers. Existing benchmarks do not test whether agents can take over this requirement–code con- sistency audit. Implementation benchmarks do the check for the agent with pre- pared tests, and consistency checkers are given the function to examine. In a real repository, the documentation states many requirements and marks none as broken, so the auditor must first decide which promise to check. We introduce ReqAuditBench to measure this proactive investigation. Each task gives an agent a repository and a region of its documentation that contains at least one violated requirement, without saying which one. We build tasks from pull requests that update both a documented requirement and its implementation. We place the up- dated documentation into the old repository, where nothing marks which passage changed. On 109 tasks with 134 requirements from 9 Python repositories, eight coding agents reach only 35.1–57.5% requirement recall, the share of violated re- quirements they report. Seeing a requirement is often not enough. For 60.3–92.6% of missed requirements, the promise had already appeared in the agent’s tool out- put. In case studies, agents stop before checking the code behind the promise. The results suggest that coding agents are not yet ready to take over this audit.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.