LOD: Latent Objective Discovery in Heterogeneous Multi-Agent Reinforcement Learning
Abstract
In multi-agent reinforcement learning (MARL) settings, agents are often heterogeneous with diverse intrinsic utility functions. In this study, we introduce **L**atent **O**bjective **D**iscovery (**LOD**) strategy among agents to decompose their distinct utilities into mixed underlying objectives in a low-dimensional latent space. Latent objective discovery aims to capture flexible and abstract coordination patterns among agents and each abstract objective is associated with a shared value function. During learning, all the learned values are aggregated by each agent according to their personalized weights learned in an adaptive manner. Therefore, agents inherit objective-level estimates for policy updates and value learning, enabling structured information sharing without requiring access to other agents’ policies. Algorithmically, we incorporate our approach into Multi-Agent Proximal Policy Optimization (MAPPO) to exploit this structure. Theoretically, we establish convergence guarantees under linear function approximation within the actor-critic framework. Empirically, we extensively validate the advantage of introducing latent objective discovery in popular multi-agent testbeds with heterogeneous agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.