acceptodds
Under review as a conference paper at ICLR 2027

Share First, Specialize Later: Temporal Parameter Sharing in Multi-Agent Reinforcement Learning

Abstract

In multi-agent reinforcement learning (MARL) behavioral diversity among agents is crucial when a task requires them to pursue distinct goals, cover different areas, or take on asymmetric roles. Existing methods induce this diversity by modifying the training objective or the policy architecture, often at the cost of extra structural and hyperparameter complexity. We instead induce diversity through training dynamics alone, via temporal parameter sharing: agents first pre-train a single shared, homogeneous policy, then split into independently trained, heterogeneous policies initialized from it. We formalize the choice of how long to pre-train homogeneously as a bilevel optimization problem. Since its non-convex lower level cannot be differentiated through, we solve it with Bayesian optimization over the switch point — an approach we call BOSS (Bayesian Optimization for Shared-to-Split training). Across four multi-agent tasks (navigation, dispersion, predator-prey, and sampling) spanning three MARL algorithms (IPPO, MADDPG, IDDPG), BOSS matches or even exceeds DiCo, a state-of-the-art diversity-control baseline evaluated at its own best published configuration without any change to the policy architecture. The variation in the switch point found by BOSS tracks how symmetric each task’s underlying coordination problem is, positioning the duration of homogeneous pre-training as an interpretable, previously underappreciated diagnostic of task structure rather than a hyperparameter to tune away.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.