Online World-Model-Boosted Planning for Multi-Agent Quadruped Soccer
Abstract
Learning coordinated behaviors in multi-agent systems remains challenging due to sparse rewards, strong inter-agent coupling, and large joint action spaces, which make model-free exploration highly sample-inefficient. We study this problem in multi-quadruped robot soccer, which is a representative and commonly-used benchmark. Specifically, each robot selects among pretrained repositioning, dribbling, and kicking skills together with continuous command parameters while competing against mobile opponents through self-play. We propose an online model-based multi-agent reinforcement learning framework in which model predictive control (MPC) method serves as a continually improving training-time teacher for decentralized high-level policies. A skill-level world model is learned online from self-play transitions and used by an MPC to optimize hybrid skill selection and continuous skill parameters. To compensate for the limited planning horizon, a learned centralized value function provides long-term guidance during action evaluation. The resulting MPC decisions are subsequently distilled into decentralized high-level policies for efficient execution. This creates a closed-loop learning process in which evolving policies generate increasingly relevant data for world-model learning, while the improving model provides progressively stronger planning targets for policy optimization. Experiments against model-free Multi-Agent Proximal Policy Optimization (MAPPO) and ablations show that the proposed method improves environment sample efficiency and overall policy performance. The project website is available at https://anonymous.4open.science/w/skills2strategy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.