Discovering Semantically Grounded and Variable-Duration Multi-Agent Cooperation Skills from Offline Multi-Task Data
Abstract
Offline multi-task MARL learns cooperative policies from static source-task datasets for zero-shot deployment to unseen tasks. Recent skill-based methods improve cross-task generalization by discovering reusable latent coordination skills. However, existing methods learn distinguishable latent skill representations without semantic guidance; this leaves the skill-discovery space overly broad, and the resulting distinctions may not correspond to semantically meaningful coordination behaviors. Moreover, step-wise or fixed-interval skill inference prevents the skill durations from adapting to changes in the team's high-level coordination intention. Accordingly, we propose Semantically Grounded and Variable-Duration skill learning (SGVD). We uses an LLM-assisted pipeline to iteratively refine Label Functions (LFs), which provide weak supervision through semantic labels and transition boundaries. Semantic labels guide latent skill representations to capture distinct and recognizable coordination behaviors, while transition boundaries supervise an observation-conditioned predictor that detects changes in the team's high-level coordination intention during execution and triggers coordination-skill transitions accordingly. Furthermore, we introduce segment-terminal prediction and dynamics distillation as two auxiliary tasks to enhance the ability of coordination-skill representations to capture segment-level transition dynamics. Experiments on offline multi-task SMAC task sets under four data qualities show that SGVD improves zero-shot performance while discovering semantically grounded, variable-duration coordination skills.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.