acceptodds
Under review as a conference paper at ICLR 2027

Endogenous Byzantine Faults: Measuring and Containing Fault Contagion in LLM Multi-Agent Systems

Abstract

LLM multi-agent systems use voting, debate, and peer review to obtain reliability from redundancy, on a premise inherited from Byzantine fault tolerance (BFT): an honest majority outvotes a faulty minority. The premise treats the set of faulty agents as fixed by the adversary. We show that in LLM systems the fault set is endogenous: an honest agent becomes faulty by reading a protocol message, so the fault spreads through ordinary use of the protocol. We formalize a contagious Byzantine fault model with three measurable quantities, the per-hop infection rate , the payload replication rate , and the inter-replica failure correlation , under which an honest majority survives fewer than adversaries under independent compromise, i.e. none once ; contagion across rounds adds a cascade condition and correlated replicas an irreducible failure floor. The vote share these quantities imply tracks debate accuracy over 62 systems (, held-out , no better than alone) and places 53 of them on the correct side of the majority threshold without fitting. In about 780k agent calls over 11 models from seven families and 1300 questions, a message that only claims coordinator authority flips 97% of Llama-3.3-70B agents on arithmetic they can verify themselves, while a worked argument flips 0%. Within the model families we test, larger models are harder to persuade but easier to command, and some models forward the payload at every hop of a relay chain. One adversary among five drives debate accuracy from 74–90% to 2% or less for four large models. Mixing vulnerable model families does not help; only members that resist the instruction channel do. A provenance gate that admits peer evidence but not peer directives restores accuracy to 72–87% and, unlike a warning prompt, mostly preserves the rate at which agents accept legitimate corrections; an adaptive attacker is left with the weaker evidence channel, which the gate does not address.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.