acceptodds
Under review as a conference paper at ICLR 2027

Social Pressure Breaks Majority Voting in LLM Safety Panels

Abstract

Large language models (LLMs) are increasingly used to judge whether content is safe. One way to reduce errors from any single model is to combine judgments from several models by majority vote. Yet LLMs are also known to change their judgments after seeing others’ answers. This raises a basic question: does majority voting still correct errors when every reviewer sees the same misleading peer message before voting, or does it instead amplify them? Across six open-weight LLMs and six safety datasets, we hold the reviewers and voting rule fixed and change only the message shown before the vote. We find a sharp reversal. Under silent peers, the panel’s false-alarm rate on benign items is 42.0%, below the 55.7% average across reviewers. Under wrong peers, reviewers falsely flag 86.5% of benign items on average, and the panel flags every benign item. This failure does not require reviewers to make the same errors: at these error rates, majority voting would fail essentially every time even if their errors were independent. But the panel can recover. Keeping three of six votes made before the peer message restores the panel’s advantage. Conformity therefore does more than change individual judgments: once those judgments are combined, the same misleading peer message can turn majority voting from correcting errors into amplifying them.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.