acceptodds
Under review as a conference paper at ICLR 2027

SWE-AaaV: Building Adaptive Executable Evidence for Coding Agent Rewards

Abstract

Large language models have made rapid progress on software engineering tasks, where pre-defined unit tests are widely used as verifiable rewards for training coding agents. However, limited coverage and implementation-specific assertions can cause them to accept incorrect solutions or reject valid ones, and such misclassifications are difficult to eliminate at scale. In this paper, we propose SWE-AaaV, an Agent-as-a-Verifier framework that constructs adaptive executable evidence for coding agent rewards. Rather than directly judging patch correctness, it generates paired test scripts conditioned on the candidate and reference solutions. To reduce false negatives, the paired scripts accommodate implementation differences while testing the same problem-scoped behavior. To uncover false positives, the verifier agent uses white-box differential testing to find issue-relevant failures and regressions, with the reference solution serving as a behavioral oracle. A deterministic harness checks reference validity, base consistency, and resolved evidence, returning feedback for test revision. For validated evidence, execution on the candidate determines the final reward. To make this candidate-conditioned verification practical for reinforcement learning (RL), we integrate it with asynchronous reward computation and hybrid scheduling. In offline evaluation, SWE-AaaV improves verifier AUC by 4.55-5.98 points over the strongest model-based baselines. When used to provide RL rewards, SWE-AaaV improves SWE-bench Verified performance by 1.8 points over standard unit-test rewards, with only a 15% increase in wall-clock training time. These results demonstrate a practical approach to evaluation-time scaling for coding agent rewards.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.