acceptodds
Under review as a conference paper at ICLR 2027

Beyond Binary Tests: Auditing Execution-Derived Rewards for Repository-Level Agent RL

Abstract

Repository-level coding agents are attractive targets for reinforcement learning because software can be executed, yet training commonly reduces that evidence to a single test-pass bit. We study two limitations of this paradigm: verifier trustworthiness, because passing tests need not imply that the intended task was solved, and reward starvation, because uniformly failing rollout groups provide no within-group learning signal. We audit execution-derived rewards for selectively certified repository tasks by separating two choices: which executable evidence the environment collects and how that evidence is aggregated. Holding hidden reference executions fixed, we characterize exactly when strict all-pass and fractional aggregation can select differently, then test semantic validity and held-out utility. Across 42 prospectively evaluated model-repair tasks, raw selection opportunity is 0/42 at the predeclared endpoint and 1/42 at and . In 288 natural agent patches, the sole raw opportunity among 18 comparable pools disappears under repository-native semantics. Controlled studies exhibit positive held-out differences on some raw-opportunity pools, but gains are concentrated; a task-disjoint replication has 2/16 raw opportunities and an average-effect interval including zero. Complementary historical real-agent data show that an earlier execution-derived signal recovers variation in 19/22 terminal-binary-flat rollout groups versus 8/22 for repository-test fraction. A frozen adversarial campaign withholds full credit from 21/22 qualifying visible-verifier exploits, while a separate optimizer-interface check verifies a nondegenerate optimizer signal without supporting a causal policy comparison. These results motivate selective executable supervision: determine whether the bottleneck lies in reward semantics, aggregation, evidence construction, or candidate generation rather than forcing every task into a binary or pseudo-dense reward template.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.