acceptodds
Under review as a conference paper at ICLR 2027

AutoApproveBench: Benchmarking LLMs as Auto Reviewers for Coding Agents

Abstract

As LLM agents gain the ability to execute long-running tasks autonomously, their actions increasingly carry risks of compromising safety or exceeding user authorization. Model-based action approval has emerged as a crucial safeguard, yet existing evaluations remain proprietary and product-specific, precluding open and controlled comparison across reviewer models. We introduce AutoApproveBench, the first open and controlled benchmark for systematically evaluating LLMs as auto-approval reviewers. AutoApproveBench curates 1,579 context-grounded approval events from 75,165 review-relevant records drawn from real coding-agent sessions, enables controlled cross-model comparison, assesses both correct approval and correct denial, and organizes all denial cases into 3 diagnostic categories covering unsafe actions, violations of explicit requirements, and unrequested scope expansion. Extensive experiments with 10 representative models on AutoApproveBench show that (I) across models, correct denial rates range from 84.58% to 95.42%, while correct approval rates range from 31.95% to 82.31%. Detailed analysis further reveals that (II) reviewers occupy substantially different error profiles between preserving authorized work and blocking inadmissible actions, and (III) these points do not monotonically track general-purpose model capability. These results establish auto-approval reviewing as a distinct capability dimension that demands dedicated evaluation and calibration. Our work is openly accessible.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.