Trust Before Autonomy: Pre-Outcome Selective Routing for Coding Agents
Abstract
Coding agents act before their final correctness is known. We study pre-outcome selective autonomy: whether an exact software change should proceed automatically, be reviewed, or be withheld using only evidence available before grading. OpenCoder X binds the candidate to uncertainty, evidence provenance, verification status, and run identity while enforcing the boundary y not in z between pre-outcome information and the eventual benchmark label. A frozen policy then maps that record to ACCEPT, REVIEW, or ABSTAIN. Across 128 labeled SWE-bench Verified candidates, two failure modes emerge. In a matched 68-candidate development pilot, Guard changes resolved-task rate from 12/17 to 13/17 for the Codex client but from 13/17 to 11/17 for the Claude Code client; paired inference is inconclusive, and the effects differ in sign. On a disjoint 60-candidate calibration split, both clients resolve 20/30 tasks, yet every candidate receives raw failure risk 0.35 (AUROC 0.50; Brier 0.2225). Isotonic calibration maps the single support point to failure probability 1/3, after which the frozen policy abstains on every candidate. The failure is therefore not merely miscalibration: the score contains no instance-level ordering to exploit. More generally, deterministic scalar calibration cannot separate candidates that are tied by the underlying score; K distinct score values yield at most K + 1 threshold-defined empirical coverage regimes. The contribution is a leakage-resistant protocol for coding-agent authorization together with an empirical diagnosis of when selective routing fails because the pre-outcome representation itself lacks discrimination.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.