Mitigating Representation Drift in MARL with Trust-region Autoencoders
Abstract
Vision-based multi-agent reinforcement learning (MARL) suffers from poor sample efficiency, limiting its practicality in real-world systems. Representation learning with auxiliary tasks can enhance efficiency; however, existing methods, including contrastive learning, often require the careful design of a similarity function and increase architectural complexity. In contrast, reconstruction-based methods that utilize autoencoders are simple and effective for representation learning, yet remain underexplored in MARL. We revisit this direction and identify unstable representation updates (representation drift) as a key challenge that limits the sample efficiency and stability in MARL. To address the challenge of representation drift, we propose the Multi-agent Trust Region Variational Autoencoder (MA-TRVAE), which stabilizes latent representations by regularizing updates within a trust region. Combined with Multi-agent Policy Proximal Optimization (MAPPO), MA-TRVAE improves its sample efficiency and stability in vision-based multi-agent control tasks. Experiments demonstrate that this simple approach is competitive with state-of-the-art MARL methods, while being more computationally efficient. Furthermore, we show that our method can scale up to more agents with only slight performance degradation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.