The Last Mile of Skill Transfer: In-Sample Team-Value Alignment for Offline Multi-Agent Generalization
Abstract
Offline multi-agent reinforcement learning (MARL) from multi-task datasets aims to learn cooperative policies that generalize zero-shot to unseen tasks with varying numbers of agents, and hierarchical skill learning is the dominant approach. In these methods, however, value supervision stops at the skill: the low-level executor is trained by conditional behavior cloning, so successful and unsuccessful realizations of one skill are blended, and a correctly selected skill is still executed by low-value actions. We formalize this as the hierarchical value gap, which closes only under a signal that varies with realization quality within a skill, and that no skill-level objective supplies. We propose Skill-conditioned Team-value Execution Projection (STEP), a training-time framework that completes the value pathway of hierarchical skill policies by extending team-value supervision from skill selection to skill execution. Our approach introduces an in-sample expectile team value that distinguishes the quality of logged executions under the same skill and supplies the weighting signal for execution projection, a team-residual execution projection that uses this signal to make decentralized executors imitate high-value realizations of each skill, and execution-consistent consequence conditioning that aligns the value with the skill semantics actually being executed. On nine unseen SMAC maps, STEP improves over the strongest hierarchical skill baseline by 5.4 and 4.8 win-rate points on medium-expert and medium-replay data, respectively, while matching its performance on expert data, whose returns are saturated. These results show that STEP delivers value supervision to skill execution, strengthening zero-shot cooperation on mixed-quality data without sacrificing performance on high-quality data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.