acceptodds
Under review as a conference paper at ICLR 2027

Better Agents, Worse Teams: When Low-Order Tests Certify Independent Upgrades

Abstract

Independently improved agents can form a worse team. We study when tests of single and paired replacements certify every allowed mixed-version deployment. Bounds on policy changes define an exactly realizable activation class. We derive its sharp pairwise remainder and prove that, for agents with equal exposure , the anchored predictor is minimax exactly when . Verified causal dependencies give smaller remainders and polynomial target bounds. Synthetic experiments compare different statistical information: on a 16-agent workflow, pairwise testing certifies three of three pools at 64 million episodes with a residual-coupling known-moment bound, and at 128 million with empirical Bernstein and common random numbers. Independently fitted 12–16-agent teams include compatible and incompatible upgrades; pilot-allocated exhaustive testing narrows some cost gaps. Multi-step local tasks reveal failures missed by uncorrected pairwise fits. Episode gates certify only the gate-averaged canary mixture; small gains require low gate rates and large evaluation budgets. Exact feasible intervals and attained-remainder sampling separate uncertainty caused by missing information from conservative statistical bounds. The results identify when low-order tests are informative, optimal, and cheaper than direct evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.