ART: Auxiliary Representation Transfer for Multi-Agent Policy Optimization
Abstract
Actor-critic methods in multi-agent reinforcement learning (MARL) typically rely solely on scalar rewards to shape actor representations. This narrow training signal leaves rich joint system dynamics unexploited, which severely limits performance and adaptability, especially under partial observability. We overcome this by introducing Auxiliary Representation Transfer (ART), a novel framework that aligns decentralized actor features with the predictive latent coordinates of a centralized Koopman autoencoder. Unlike prior methods that feed global models directly into the policy, ART regresses these coordinates through a linear auxiliary head during training to embed global predictive dynamics directly into the action selection process. Crucially, this transfer requires no execution-time overhead and leaves the actors' local observations and network architectures entirely unchanged. Experiments on Multi-Agent MuJoCo demonstrate that ART substantially improves the sample efficiency and final performance of Multi-Agent Proximal Policy Optimization (MAPPO). Notably, ART yields up to a 76% performance gain on Ant under full observation and more than doubles the final return on Ant under partial observation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.