Which Teams Must Be Evaluated? Cost and Robustness Limits of Compositional Agent Evaluation
Abstract
Evaluating every combination of frozen agents can be costly. We study simultaneous value estimation with charges for team preparation and each evaluation trial, conditional on a supplied valid bound on mean-model error. Under fixed Hoeffding certification, we identify exactly when sharing observations cannot reduce any team's direct sample requirement. A tensor identity computes this boundary for complete bipartite catalogues. For a four-role path, the cheapest plan shifts from sixteen teams to twelve with unequal sampling, then eight, as preparation prices rise. With model error, preparing eight rather than sixteen teams can require 39 times as many observations. Positive observation prices preserve the boundary but change the preferred allocation. A finite-pool evaluator compares empirical-Bernstein and direct confidence-sequence procedures, charges selection costs, and confirms the selected procedure on independent observations. Controlled studies distinguish structural sharing, support adaptation, and procedure choice. These results characterize acquisition-cost trade-offs under stated model and certificate assumptions; valid confirmation alone does not guarantee cost-optimal selection or deployment savings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.