acceptodds
Under review as a conference paper at ICLR 2027

Beyond Mimicry: A Value-Based Approach to Task-Efficient Offline MARL

Abstract

Offline Multi-Agent Reinforcement Learning (MARL) offers a vital path to complex coordination without the risks of real-world interaction. However, most existing methods are task-specific, requiring retraining for new tasks, leading to redundancy and inefficiency. To address this issue, we propose a task-efficient value-based multi-task offline MARL algorithm, Task-efficient Skill-based Q-learning (TasQ). Unlike existing methods that rely on behavior cloning for skill execution, TasQ adopts a value-centric framework. It discovers skills in a latent space by reconstructing the next observation, encouraging the extraction of task-agnostic features that transfer across tasks. It then evaluates fixed and variable actions separately, and uses conservative Q-learning with local value calibration to select the optimal action for each skill, enabling strong multi-task generalization from limited source tasks. Substantial experiments spanning homogeneous and heterogeneous agents, single and multiple source tasks, and discrete and continuous actions demonstrate the superior generalization and task-efficiency of TasQ. It achieves the best performance on out of task sets, with up to **68.9%** improvement on individual task sets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.