Conformity Breaks Conformal Prediction
Abstract
Multi-agent LLM systems often rely on models that answer after seeing peer messages. Conformal prediction can help decide when an answer is reliable enough to act on, but calibration is usually done with the model answering alone. We identify and formalize a score-mechanism shift: the question stays the same, but peer messages change how the model scores its answers. We show that this one shift causes failures at three levels. First, peer pressure breaks overall coverage. Across four open-weight LLMs in our core experiments, coverage falls from a calibrated 90% to 74% under unanimous-wrong peer messages. More importantly, the average hides a much larger failure. On questions the model is already less sure about, coverage falls from 87% to 47%. An attacker can also identify such weak cases without knowing the correct answers, using only outputs from the target model or predictions from another model. Finally, the problem changes what the system actually does. A policy designed to defer uncertain cases instead acts on the wrong answer in 12% of cases where peer pressure flips the model's top choice. The coverage failure also holds across additional models and seven task types. Recalibrating on data with peer messages restores coverage, but makes the system defer more often. More broadly, conformal guarantees for multi-agent LLMs must account for the peer messages a model sees when it answers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.