Evidence Under Pressure: The Peer-Override Effect in Collaborative LLM Reasoning
Abstract
Large language models increasingly operate in multi-agent systems where agents exchange not only evidence but also conclusions. Yet when a peer opinion conflicts with the evidence available to a model, it is unclear whether the model preserves its own evidence-grounded reasoning. We introduce a controlled benchmark that isolates this question by holding the rules, query, and model-visible evidence fixed while varying the presence of a conflicting peer opinion. Across 500 cases and 14 LLMs, conflicting peer opinions reduce accuracy by 30.45 percentage points. Among cases answered correctly when sufficient evidence is presented alone, 53.51% are answered incorrectly when the same evidence is accompanied by a conflicting peer opinion. In contrast, only 4.31% of all cases shift from incorrect to correct. We call this systematic displacement the Peer Override Effect. Its magnitude varies by more than 55 percentage points across models with similar reasoning accuracy, revealing a substantial gap between reasoning competence and collaborative reliability. The effect persists across surface realizations and even after a model has already produced a correct answer, while explicit evidence re-evaluation only partially mitigates it. Reliable multi-agent systems must therefore not only reach evidence-supported conclusions, but also preserve them when peer opinions conflict with the available evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.