What Gets Deployed? Deployment Accounting for Stateful Agents
Abstract
Agentic RL studies usually report single-run success or oracle pass@k, yet practical systems may generate several sandboxed executions and release only one. These metrics can therefore misidentify both the best system and the reason it improves. We introduce deployment accounting: a protocol that jointly reports native success, successful commitment, wrong commitment, abstention, and oracle availability, and that decomposes policy differences into candidate quality and the selector's incremental value on identical pools. We instantiate a label-independent outcome selector and conduct a controlled 7B-scale study on ALFWorld and a restricted WebShop setup. Selection adds roughly 9–13 success points to the same ALFWorld weights across the evaluated learned policies. Over four continuation seeds, a simple penalty-free outcome-RL control reaches 72.5% committed success in seen rooms and 71.3% in unseen rooms, compared with 62.9% and 60.5% for the original GRPO pipeline. A commit-aware variant reaches 79.2% in the designated unseen-room run and provides favorable low-error operating points, but a matched zero-credit control shows that its auxiliary credit is not required for the strongest average policy. Same-pool analyses attribute most average gains to better candidates rather than a larger voting increment; WebShop shows that fewer wrong commitments can instead come from more abstention. These results establish deployment accounting as a practical methodology for separating learning, selection, and coverage effects before attributing improvements in stateful-agent systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.