acceptodds
Under review as a conference paper at ICLR 2027

PROWL-2: Fidelity-Gated Exploration for Multi-Agent World Model Dual Curricula

Abstract

World models promise sample-efficient multi-agent reinforcement learning by letting policies learn from imagined interaction. Imagination, however, introduces a fundamental ambiguity: a trajectory with high learning signal may expose a genuine weakness of the task-solving policy, or may simply reflect inaccurate world-model predictions. A curriculum that prioritises such experience without accounting for model fidelity can therefore reinforce model errors rather than improve the agents. We introduce PROWL-2, a fidelity-gated dual-curriculum framework that maintains one curriculum for training the task agents and a second for discovering and repairing world-model failures. PROWL-2 scores imagined experience along two axes—learning potential and model fidelity measured against the real environment—and routes it accordingly: reliable, informative rollouts are promoted to the agent curriculum, while unreliable ones are sent for world-model repair. A dedicated developer policy actively searches for further model failures, and repaired regions are re-evaluated before re-entering agent training, closing the loop between world-model improvement and policy learning. On cooperative tasks in StarCraft Multi-Agent Challenge v2 (SMACv2) and the Multi-Agent Quadruped Environment, spanning discrete combat and continuous robotic collaboration, PROWL-2 consistently outperforms strong model-based and curriculum baselines at matched environment interactions. Our results show that learning inside world models requires deciding not only which imagined experience is useful but which is trustworthy and that making this distinction explicit turns world-model failures into a source of policy improvement rather than a cause of policy error.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.