LLM Cognitive Bias Exploitation and Emergent Consensus Capture in Connected Agents
Abstract
LLM agents increasingly deliberate in connected panels, where every agent sees the verdicts of the others. This turns a cognitive bias of the models themselves, their tendency to defer to peers, into an attack surface: an adversary who can post one verdict may pull the honest agents with it. We ask whether that failure can be predicted before deployment from a measurement on a single agent. It can, in closed form. If an agent abandons its own evidence once a unanimous peer margin reaches , then public deliberation acquires regularities that belong to no individual agent, including an error floor of , where is the accuracy of an agent's own evidence and , that does not fall as the panel grows and capture of a panel of any size by early adversaries. Both and are measurable, so the prediction has no fitted parameters. We test it over 4,770 panel episodes on OpenAI and Anthropic models. Measured with the classic Asch item, where an agent commits and is then contradicted, the prediction fails badly (calibration slope 0.300, where 1 is perfect, against a pre-specified ). Measured the way deliberation actually presents peers, before the agent commits, the same prediction is calibrated (slope 0.827, ). Conformity recovered from 4,980 decisions inside the panels shows why: real agents behave as the in-context measurement says, not as the Asch one does, and they defer gradually rather than at a threshold. Biases also compound, but selectively. An adversary that adds one unverifiable claim of superior evidence to its verdict doubles the damage and captures panels of the most deferential configuration 93% to 97% of the time, whereas claims of rank or of certainty add almost nothing: these agents are moved by what resembles evidence and not by status. The same asymmetry supplies defenses. Dissent that merely disagrees is largely ignored, while the identical dissent carrying an evidence claim halves mean error, and cuts it from 0.933 to 0.233 for GPT-4.1-mini; sealed voting removes the mechanism by construction; and panels of reasoning models largely resist the attack. Code, items, and all 61,706 cached model responses are in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.