SAW-v2: Supervised Dual-Stream State Adaptive Weight for High-Dimensional Multi-Agent Reinforcement Learning
Abstract
Multi-agent reinforcement learning (MARL) has advanced significantly in complex robotic systems, enabling decentralized agents to acquire coordinated behaviors for locomotion, dexterous manipulation, and cooperative decision-making under the centralized training and decentralized execution (CTDE). However, existing MARL algorithms often process heterogeneous and high-dimensional observations in a largely uniform manner, which can dilute task relevant state information, overemphasize redundant or noisy features, and make actor-critic optimization less efficient and less stable in strongly coupled multi-agent environments. To address this limitation, we propose SAW-v2, a state adaptive weighting framework that learns stable state-importance weights for actor side representations. By identifying informative observation components and suppressing less useful ones, SAW-v2 provides a lightweight plug-in module for improving policy learning. 108 groups experiment on MAMuJoCo, DexHands, SMACv2, and MPE demonstrate that SAW-v2 consistently improves representative MARL baselines across both continuous control and discrete cooperative benchmarks, achieving average 12.3% performance gains and showing stronger robustness in heterogeneous and high-dimensional coordination tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.