Data, Objective, and Architecture: When Do Independent Policy-Gradient Learners Collude?
Abstract
Generative AI agents increasingly make sequential decisions in multi-agent environments, where independently reward-maximizing agents can nevertheless learn cooperative or collusive behavior. We study how two defining components of modern generative agents— and —shape collusion under independent policy-gradient learning. Using a repeated pricing game as a controlled social-dilemma environment, we develop an analytical framework that separates what an agent learns its pre-training data, what it learns under its post-training objective, and what it can through its context architecture. We first show that next-token pre-training acts as a behavioral-transmission mechanism: under realizability and standard generalization conditions, the learned policy recovers the teacher's history-conditional action distribution, so pre-training determines the behavioral initialization from which subsequent multi-agent learning begins. We then characterize independent own-reward population policy-gradient dynamics. With one-step memory, we derive an exact action-value gap showing that the gradient favors high-price actions only when the opponent's learned reciprocity exceeds an explicit threshold. This yields a patience condition and an invariant non-collusive basin: sufficiently non-collusive initial policies converge toward competitive behavior, while crossing the threshold is not by itself sufficient for collusion. Memory changes these dynamics by creating additional continuation-value channels. Without learned memory, agents cannot retaliate, but collusion can still arise through coordinated discovery; we derive a separate threshold below which policy-gradient learning eliminates high-price behavior. For learned context length , we analyze a soft-attention policy and derive a sufficient non-collusion threshold that decreases with while remaining strictly positive for every fixed discount factor below one. Thus, longer memory enlarges the set of future states through which current actions can affect an opponent's behavior and correspondingly shrinks the guaranteed competitive basin. Together, our results show that pre-training shapes multi-agent learning through , while memory reshapes its , providing a tractable learning-theoretic account of emergent collusion under independent policy-gradient learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.