acceptodds
Under review as a conference paper at ICLR 2027

On the Effect of Shared Memory on False-Belief Lock-In in Multi-Agent Systems

Abstract

When AI agents share a written memory, does the group become more reliable, or more gullible? A shared record can preserve useful information, but it can also spread a false claim so that agents gradually become entrenched in a false belief, which we call false-belief lock-in. We simulate a persistent source of error with a compromised agent by instructing some agents to repeatedly endorse a false claim, and ask the remaining agents to judge the claim for themselves. In controlled comparisons with 6- and 12-agent groups, the remaining agents endorse the false claim when they communicate through a shared record, but not when they speak directly to one another or see only their own earlier judgments, even when the repeating agents endorse the claim on every turn. Even a small minority can spread the claim to the group through a shared record. In groups of twelve, two agents lead the other ten to endorse it in 70% of their final judgments. In groups of one hundred, it can spread to every agent judging for themselves, although whether it does depends strongly on speaking order. Corrections and warnings that repeated entries are not independent evidence can reduce the endorsement of false claims in several settings; however, no change works in every setting. Our results demonstrate false-belief lock-in as a plausible safety concern given compromised agents and shared memory.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.