acceptodds
Under review as a conference paper at ICLR 2027

Balanced Multi-Agent Reinforcement Learning via Desirability-Gated Pseudo-Expert Regularization

Abstract

The exploitation–exploration trade-off remains a central challenge in complex cooperative multi-agent tasks, where agents must discover coordinated joint behaviors from local observations. Entropy- or divergence-regularized actor-critic methods partly alleviate this challenge, but they often provide only weak guidance for learning complex team behaviors. Exploiting expert behaviors via behavior cloning can reproduce such behaviors, yet it usually assumes access to expert demonstrations and is therefore difficult to apply in online settings. To address this issue, we propose Balanced Multi-agent Reinforcement Learning (BMARL), a desirability-gated policy-regularization framework that balances exploration and exploitation using a pseudo-expert learned online from desirable trajectories. BMARL introduces a hybrid divergence-regularized objective: on desirable trajectories, the learning policy is regularized toward the pseudo-expert prior, while on undesirable trajectories, the objective reverts to maximum-entropy exploration. The framework is compatible with standard MARL backbones and mixer choices, and we provide theoretical analyses of policy evaluation and improvement. The proposed method is evaluated on various MARL tasks, including dense- and sparse-reward settings, and achieves further performance improvements over state-of-the-art methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.