Team Calibration for Multi-Agent Imitation under Decentralized Information
Abstract
Multi-agent imitation systems must turn centralized demonstrations into high-value decentralized behavior, because teacher information and team objectives are not preserved by separate local decisions. Overcoming this mismatch requires calibration to executable team utility, rather than accurate reproduction of agentwise action labels. Local imitation is well suited to decentralized deployment, yet its reliance on teacher marginals leaves the utility-optimal team protocol unidentified when value depends on joint behavior, especially in symmetric teams with private signals. Within that setting, joint-occupancy imitation captures demonstrated coordination through distribution matching, but its target remains tied to demonstrated behavior and cannot resolve utility interventions that preserve the demonstrations. We therefore propose Structured Team-Calibrated Projection (\stcp), which constructs a fixed executable portfolio by applying public-context routing and shared-clock switching to theory-identified decentralized atoms and selects its deployed member by held-out team value; our analysis provides an exact phase characterization, a heterogeneous convex-closure result, and a finite-portfolio selection bound. Experiments on controlled utility interventions, synchronization tasks, and graph consensus show that recovers the predicted protocol transitions and improves team value over representative decentralized baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.