Group Alignment: Internal and External Incentives for Cooperation in Social Dilemmas
Abstract
AI agents are becoming increasingly autonomous and, similarly, act in more and more complex environments. In such environments, agents pursue underspecified goals and need coordinated interaction with other agents, however, how this cooperation is undertaken is still poorly understood. In this paper, we characterize the cooperation between different agents in scenarios and prove an inefficiency theorem of contracts in such complex environments. We show two orthogonal categories of solutions to the collaboration problem grounded in different economic theories of cooperation, namely contracts and "we”-frames. Frames internally shift reasoning patterns of different agents towards cooperative behaviour and enable finding the best individual mutually beneficial action for the group, while contracts align incentives with an external device for the same objective. We study the effect of a range of contract representations, from natural language to formal contracts across all frames in classical matrix-based games and common-pool resource environments. We find that in matrix games, contracts reliably improve mutual-advantageous actions. We observe that contracts of any kind perform 15.33% worse in environments where agents choose selfishly compared to when they instantiate the group frame, and 41.76% worse in common-pool resource environments. Our results suggest both self-negotiated contracts and "we”-frames improve cooperation over LLM behaviour without such devices, with "we”-frames being more robust in complex settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.