acceptodds
Under review as a conference paper at ICLR 2027

Test-Driven Reward Design for Cooperative Multi-Agent Reinforcement Learning

Abstract

In reinforcement learning (RL), especially in multi-agent reinforcement learning (MARL), reward design remains a fundamental challenge, as dense rewards often require task-specific engineering, while sparse rewards provide insufficient learning guidance for effective coordination. We observe that task completion can be characterized by whether agents satisfy task-level behavioral requirements, which can serve as implicit tests for evaluating cooperative behaviors. Based on this insight, we adapt the paradigm of Test-Driven Reinforcement Learning (TdRL) to the multi-agent setting and develop **T**est-**D**riven **M**ulti-**A**gent **R**einforcement **L**earning (TdMARL), a framework that constructs return functions from task-level evaluations and derives informative reward signals for the multi-agent system, enabling a more intuitive and systematic design process for complex cooperative environments. We provide theoretical analysis establishing the validity of extending the test-driven paradigm from single-agent RL to cooperative MARL. Furthermore, we introduce MARL-specific adaptations to TdRL, including tailored data sampling and preference-evaluation mechanisms, together with a soft uniform prior for stable return-to-reward decomposition, to better accommodate the sensitivity and non-stationarity of multi-agent systems. Experiments across dense reward, sparse reward, and sparse reward with auxiliary signals demonstrate that TdMARL provides effective reward guidance and achieves robust coordination performance across diverse reward settings. TdMARL offers a new perspective on reward design for cooperative MARL by shifting the focus from manually engineering scalar rewards to specifying task-level behavioral requirements.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.