Two-Level Bandits for Collaborator Selection and Sustained Cooperation
Abstract
Agentic systems increasingly require one autonomous agent to choose another to work with: which tool-using agent to delegate a multi-turn task to, which client to admit to a federation, which service to contract with for a sustained workflow. Such collaborator selection problems have a two-level structure that is usually studied one level at a time. At the outer level an agent searches a set of uncertain candidates under an interaction budget or deadline (a secretary / best-arm-identification prob- lem); at the inner level, once a candidate is engaged, value is produced only if the two agents sustain cooperation in a repeated game with a defection temptation and noisy execution (an iterated Prisoner’s Dilemma). We formalize the joint problem as a two-level bandit in which the reward of each outer arm is itself the equilibrium payoff of an inner repeated game, give algorithms for it, and complete the associ- ated discounted commitment-stopping problem left open in prior work. We prove three results—an inner cooperation-sustainability threshold γ⋆ = (T −R)/(T −P ), a value-decomposition regret bound showing that ranking candidates by advertised compatibility alone is suboptimal by exactly the cooperation gap, and the optimality of a monotone reservation threshold for discounted commitment—and validate all of them with a reproducible simulation suite. Two-level best-arm identification recovers the truly best collaborator with probability 1.00 at confidence δ=0.05, whereas a compatibility-only selector is captured by a high-capability defector and suffers a persistent value regret of 1.60; critically, its probability of correct selection degrades to zero as its interaction budget grows, because more data only sharpens its estimate of the wrong quantity. Under a costly clock, the textbook 1/e rule captures only 2.2 of the 22.2 units achieved by the optimal reservation policy. Our framework yields concrete design guidance for agent marketplaces, delegation layers, and federated participant selection, and a clean bridge between optimal stopping, bandits, and the evolution of cooperation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.