CoopShift: Post-Training Amplifies the Memory Curse in LLM Cooperation
Abstract
How does post-training change cooperation during repeated interaction? Within each of seven released model lineages, we compare pretrained (Base), supervised fine-tuned (SFT), and subsequent preference- or reward-optimized (Aligned) checkpoints across seven social dilemmas. Copies of each checkpoint play one another for 200 rounds, with visible history limited to the last two or forty rounds. The memory curse refers to lower cooperation when agents can see more past rounds. Our central finding is that post-training amplifies this decline. Averaged across the six binary-action games, expanding visible history reduces cooperation at every SFT and Aligned checkpoint. At each post-trained stage, these declines exceed Base's in six of seven lineages. Most of this average amplification is already present at SFT. With long history, both post-trained stages also cooperate less than Base on average in each binary-action game. In the symmetric binary games, their fitted cooperation rates fall more steeply on average than Base's as per-round cooperation costs rise, benefits fall, or noncooperation becomes more rewarding. The difference in cooperation rates between Aligned and SFT also changes from one-shot decisions to repeated play. When given the same past records, their choices depend on whether the instructions describe a final round or possible future interaction. Post-training evaluation should therefore go beyond isolated choices to assess whether cooperation is sustained in repeated interaction and how it depends on payoffs and visible history.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.