Sketch-Informed Capability Discovery in Large Agent Collectives
Abstract
Open collectives of foundation models, agents, and web services now contain thousands of entities, and a coordinator who needs a qualified entity for a task cannot afford to inspect them all. We study the coordinator's discovery problem under a two-tier query model, where each of the entity-groups publishes a free summary of its member's capabilities, alongside a costly query that reveals a random entity from a chosen group. Our main result is a threshold on the summary dimension with resolution . When the dimension is , summaries are informative enough to guide the search. Our proposed SIEVE algorithm uses the summaries to eliminate groups and then queries the remaining groups' members. With probability , SIEVE discovers and returns an entity that can perform the task and ranks in the top fraction of the collective, counted over the groups that contain the capability, with at most queries. This query complexity depends on neither nor the size of the collective, and no algorithm with a worst-case query limit can improve it by more than a constant factor. Below the dimension threshold, we prove in a Gaussian model of the summary readout that the query complexity of any coordinator that reads the summaries through inner products grows polynomially in . Experiments on synthetic populations with up to one million entities, as well as on RapidAPI and the Open LLM Leaderboard, support the predicted dimension threshold and query savings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.