acceptodds
Under review as a conference paper at ICLR 2027

What Does the Mediator Learn? Mechanistic Analysis of Learned Intervention Policies in LLM Games

Abstract

When LLM agents are applied in multi-agent scenarios, such as automated negotiation and collaborative coding, they often fail to effectively coordinate. In this work, we study a third-party mediator agent trained with reinforcement learning (RL). Rather than applying fixed rules, a learned mediator can adaptively decide when and what to communicate, improving cooperation without accessing agent internals. We evaluate such learned mediators across three LLM families (Llama, Granite, and Qwen) on four games spanning social dilemmas and pricing competition, and provide a mechanistic account of the results. Our experiments demonstrate that a mediator trained on one LLM family transfers zero-shot to others, provided the target model responds monotonically to cooperation messages. A contrarian agent that inverts this response collapses cooperation despite active mediation. Moreover, we find that the value of interventions decomposes in terms of timing, which is model-agnostic and learned by the mediator, and content, reflecting the combined effect of the mediator's messages and peer communication. Interestingly, content adds value only when the population's baseline cooperation rate falls below an empirically estimable threshold . Above that threshold, timing alone suffices. An extended multi-way controller improves cooperation for susceptible populations by learning state-dependent template selection. No single message type dominates overall, suggesting the gain comes from matching message type to game state rather than from any single dominant framing. Ablations without player communication estimate this threshold empirically, requiring no retraining or access to model internals. These ablations also provide a practical criterion for predicting whether richer mediator vocabularies will help in a new deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.