Beyond Mimicry: A Value-Based Approach to Task-Efficient Offline MARL
Abstract
Offline Multi-Agent Reinforcement Learning (MARL) offers a vital path to complex coordination without the risks of real-world interaction. However, most existing methods are task-specific, requiring retraining for new tasks, leading to redundancy and inefficiency. To address this issue, we propose a task-efficient value-based multi-task offline MARL algorithm, Task-efficient Skill-based Q-learning (TasQ). Unlike existing methods that rely on behavior cloning for skill execution, TasQ adopts a value-centric framework. It discovers skills in a latent space by reconstructing the next observation, encouraging the extraction of task-agnostic features that transfer across tasks. It then evaluates fixed and variable actions separately, and uses conservative Q-learning with local value calibration to select the optimal action for each skill, enabling strong multi-task generalization from limited source tasks. Substantial experiments spanning homogeneous and heterogeneous agents, single and multiple source tasks, and discrete and continuous actions demonstrate the superior generalization and task-efficiency of TasQ. It achieves the best performance on out of task sets, with up to **68.9%** improvement on individual task sets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.