SEED: Separating Admission from Deployment Evidence in Adaptive Coding-Agent Updates
Abstract
Adaptive coding agents select prompt and workflow updates using evaluator feedback. Once a hidden split helps choose the surviving update, however, it no longer supplies independent deployment evidence. We introduce SEED, a four-stage protocol that separates visible ranking, hidden admission, disjoint canary confirmation, and deployment or rollback. Partitions are fixed within each run, not resampled across generations. The analysis spans 24 model lines, six benchmark families, two serving stacks, and all 432 model–benchmark–profile cells, while distinguishing retry-inclusive gate exposures from distinct incumbent–candidate decisions. Canary checks reject 25.1%, 25.3%, and 33.6% of first admissions in the Anchor, Horizon, and Confirmatory profiles. A seemingly large retry-inclusive Anchor–Horizon increase of percentage points disappears after keeping only the first exact pair per run ( points, 95% interval ), so the raw contrast measures operational gate burden rather than increased distinct-decision non-transfer. Among 652 first-exact rollback decisions, 271 prevent a local untouched-final pass-or-gap regression and 159 prevent direct pass loss, but recall is about 33% and fixed-budget endpoint intervals all cross zero. We provide a bounded, retry-aware deployment-trace protocol and an empirical account of its behavior. The results do not show that a canary is an oracle, identify an optimal budget split, or establish contamination-free or non-coding transfer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.