Test-Time Policy Improvement for Coordination with Unseen Partners
Abstract
Cooperative multi-agent reinforcement learning prepares an agent for unseen partners by training it against a diverse population, so that one policy, fixed before deployment, generalises to partners it never trained with. But a policy committed before the partner is met must compromise whenever plausible partners call for incompatible responses, a gap no training-time diversity can close. We address this gap by introducing test-time best response (TTBR): after a short observation phase (16 episodes per seat), the agent forms an exact posterior over an explicit class of partner programs and policies and improves its own policy against it in the simulator. We show that the regret of the exact posterior response is bounded by the conflict among candidate partners, the cross-best-response regret that underlies minimum coverage sets, times the residual posterior odds, a guarantee that also predicts when partner identification cannot help. The resulting algorithm (i) responds to the partner actually present rather than to a population, (ii) retains its in-class skill, and (iii) applies to any self-play policy and pays its compute once. On Overcooked and level-based foraging, against held-out partners from the modelled family, TTBR recovers a large share of the coordination lost with strangers and beats a generalist trained on the whole partner family where responses conflict. Where nothing conflicts a generalist suffices, and outside the modelled family TTBR degrades gracefully but has no edge, while planning keeps one.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.