Executable Feedback Is a Measurement: Budgeted Acquisition and Selection-Corrected Credit for Coding Agents
Abstract
Reinforcement learning for coding agents depends on executable feedback that is costly to obtain and unreliable once obtained. Branching methods reduce cost by observing counterfactual continuations only at chosen states, but the measurements they buy are then a selected sample of the eligible credit units, and the verifier at each leaf can be wrong: on a uniform audit of validation leaves, the raw verifier disagrees with a higher-fidelity reference on 12.1% of leaves with determinate reference verdicts. We introduce BANC, a training-time layer that treats executable feedback as a budgeted, randomized measurement process. BANC purchases branch and audit measurements by their expected effect on the policy gradient against shadow-priced resource costs, calibrates raw verdicts with a selection-corrected reliability model, and recovers action-level credit with an augmented inverse-propensity estimator. Because every purchase probability is randomized, bounded below, and logged, adaptive acquisition changes which measurements are bought but not the credit estimand. Tested separately, the three mechanisms cut normalized gradient error from 0.50 to 0.34, reduce harmful update mass under 20% verifier noise from 24.7% to 9.1%, and reduce the selection bias of adaptive acquisition from 0.148 to 0.021 at 44% lower RMSE than inverse propensity scoring. Under matched token and sandbox caps, BANC improves resolved rate on a temporally held-out SWE-rebench suite by 2.7 points over BPO+Diff (p = 0.019) while using 7.4% less sandbox time than APPO, with gains concentrated on long, costly tasks and preserved on 16B and 32B backbones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.