When Should Agents Differ? Learning to Schedule Behavioral Diversity from Training Dynamics in Multi-Agent RL
Abstract
Cooperative multi-agent reinforcement learning relies on parameter sharing for sample efficiency, and a family of diversity-promoting methods counteracts the resulting behavioral homogenization. All of these methods, however, apply their diversity signal at a strength that is fixed before training begins, leaving unanswered a question that logically precedes how to promote diversity: when should diversity be promoted, and how strongly, as training unfolds? We show empirically that this temporal dimension matters: on the same task, the diversity strength that helps early in training can actively harm late-stage coordination, and no constant strength is best across all of training. We formalize diversity scheduling as a meta-level sequential decision problem over training states, and derive a tractable solution, TIDE (Training-dynamics Informed Diversity schEduler). TIDE evaluates candidate diversity strengths offline via counterfactual branching from training checkpoints, distills the resulting preferences into a 1,300-parameter gate network, and deploys the gate online to reschedule the diversity strength from inexpensive training-log statistics at deployment-time overhead below 0.01% (one-time offline cost ≈175 GPU-h per algorithm–map pair, comparable to a coefficient sweep). Wrapped around three diversity-promoting algorithms (CIA, CDS, and LIIR), each with its own features, TIDE improves the base algorithm in all 18 algorithm-map combinations on six hard and super-hard SMAC maps, outperforming hand-designed decay schedules and an online bandit. The gate’s decisions admit post-hoc interpretation, recovering task-specific diversity schedules that align with each map’s tactical structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.