acceptodds
Under review as a conference paper at ICLR 2027

Emergence and Turnover of Hidden Conventions in Monitored LLM Communication

Abstract

As LLM agents are deployed to interact with one another, their communication is often monitored to prevent collusion. Empirical studies show that monitored agents can nevertheless coordinate through hidden codes, but how such codes arise from the behavior of individual LLMs, and why they persist, is not understood. We build on and extend the theory of social conventions into a model of monitored communication, in which the agents and the monitor all learn in-context from the interaction history. The model predicts a learning race: hidden conventions let a sender pass private information once the receiver learns them, but are replaced once the monitor catches up and penalizes their use. Penalty strength selects the regime: direct communication under weak penalties, turnover of hidden conventions at intermediate ones when a convention can be established before the monitor detects it, and silence under strong ones; silence, however, stops communication only if it cannot itself serve as a hidden code. Controlled multi-agent LLM experiments in AI trading and tutoring reproduce the predicted turnover and the effects of penalty strength and monitoring lag. Causal interventions show that turnover is driven by a monitor that learns, not by penalties alone. Stronger penalties reduce flag rates while the private information keeps reaching the receiver. Monitoring thus changes how private information is shared, not whether, so a falling flag rate can mask continued leakage.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.