RoboMARL-Bench: How Task Structure and Data Composition Matter in Offline Multi-Agent Robotic Manipulation
Abstract
Offline multi-agent reinforcement learning (OMARL) requires agents to learn coordinated behavior from a fixed dataset. Yet task performance and overall dataset quality alone do not reveal how the mix of coordination experiences in the dataset affects what agents can learn. We introduce RoboMARL-Bench, a robotics-oriented benchmark with 65 task configurations, 7 rule-defined interaction structures, and 8 controlled data constructions spanning the source, episode, role, temporal, and skill layers. This design enables matched-budget comparisons within the same task while varying how experience is collected and organized. Across extensive evaluations, we find that algorithm rankings change across interaction structures, datasets with the same nominal quality can yield substantially different outcomes, and role assignment and controller ordering introduce additional task-dependent variation. Current OMARL methods also leave clear room for improvement: on data collected by suboptimal policies, they exceed the recorded data level on only about half of the tasks, and most often fall short on multi-goal and transition tasks. RoboMARL-Bench provides a controlled testbed for studying how interaction structure and offline data composition jointly shape multi-agent learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.