acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Representation Drift in MARL with Trust-region Autoencoders

Abstract

Vision-based multi-agent reinforcement learning (MARL) suffers from poor sample efficiency, limiting its practicality in real-world systems. Representation learning with auxiliary tasks can enhance efficiency; however, existing methods, including contrastive learning, often require the careful design of a similarity function and increase architectural complexity. In contrast, reconstruction-based methods that utilize autoencoders are simple and effective for representation learning, yet remain underexplored in MARL. We revisit this direction and identify unstable representation updates (representation drift) as a key challenge that limits the sample efficiency and stability in MARL. To address the challenge of representation drift, we propose the Multi-agent Trust Region Variational Autoencoder (MA-TRVAE), which stabilizes latent representations by regularizing updates within a trust region. Combined with Multi-agent Policy Proximal Optimization (MAPPO), MA-TRVAE improves its sample efficiency and stability in vision-based multi-agent control tasks. Experiments demonstrate that this simple approach is competitive with state-of-the-art MARL methods, while being more computationally efficient. Furthermore, we show that our method can scale up to more agents with only slight performance degradation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.