acceptodds
Under review as a conference paper at ICLR 2027

Cooperative agents specialize more but explore less

Abstract

Cooperation is essential for multi-agent systems to succeed. Prior work has focused on making agents more cooperative, but we argue that cooperation is only the beginning of the problem. We study deep reinforcement learning and large language model agents in sequential social dilemmas, where agents face a tension between producing immediate rewards and maintaining shared resources that sustain future production. We find cooperative agents specialize more but explore less, which can eventually undermine group performance. Otherwise identical agents who work to maximize collective rewards are more cooperative and achieve higher performance by self-organizing into maintenance or production roles. Once agents specialize, maintenance agents explore significantly less than production agents, and their search narrows further over time. They then fail to replenish sufficient resources, leaving production agents with fewer resources to use. Group performance can consequently decline by up to 93% long after cooperation has been established. These results identify a potential failure mode that can emerge after successful cooperation and highlight the importance of preserving exploration as agents specialize in multi-agent systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.