acceptodds
Under review as a conference paper at ICLR 2027

Out of Character: Per-Agent Conformal Screening for Byzantine LLM Committees

Abstract

Committees of language-model agents rely on redundancy to catch reasoning errors, but correlated exploits such as prompt injections can subvert multiple models simultaneously; when corrupted agents gain a majority, standard majority voting collapses from an uncorrupted accuracy of to . Without labels, no confidence statistic of a single decision exposes an adversary that picks a wrong action while preserving the agent's confidence profile, and peer-consensus defences fail because the adversary forms the majority. We resolve this dilemma with Sentinel, a peer-free screen that tests each agent's confidence and action choices over a window against its own clean history, bounding the probability of falsely rejecting an honest agent by a budget in finite samples under exchangeability, regardless of how many peers are corrupted. Across 5 models and 3 benchmarks, Sentinel recovers accuracy at 60% corruption and at 80% within the screened window; on free-form answers under real prompt injection it regains – of the lost accuracy; and it restores live task completion from to of tasks under a destructive attack; live under the exact null, honest rejection is – per screen with no attack, against a budget. We further prove that an adversary preserving its own action law can change the action symbol on only a bounded fraction of decisions — a bound on symbol alteration, not on task failure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.