Emergence and Turnover of Hidden Conventions in Monitored LLM Communication
Abstract
As LLM agents are deployed to interact with one another, their communication is often monitored to prevent collusion. Empirical studies show that monitored agents can nevertheless coordinate through hidden codes, but how such codes arise from the behavior of individual LLMs, and why they persist, is not understood. We build on and extend the theory of social conventions into a model of monitored communication, in which the agents and the monitor all learn in-context from the interaction history. The model predicts a learning race: hidden conventions let a sender pass private information once the receiver learns them, but are replaced once the monitor catches up and penalizes their use. Penalty strength selects the regime: direct communication under weak penalties, turnover of hidden conventions at intermediate ones when a convention can be established before the monitor detects it, and silence under strong ones; silence, however, stops communication only if it cannot itself serve as a hidden code. Controlled multi-agent LLM experiments in AI trading and tutoring reproduce the predicted turnover and the effects of penalty strength and monitoring lag. Causal interventions show that turnover is driven by a monitor that learns, not by penalties alone. Stronger penalties reduce flag rates while the private information keeps reaching the receiver. Monitoring thus changes how private information is shared, not whether, so a falling flag rate can mask continued leakage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.