acceptodds
Under review as a conference paper at ICLR 2027

Multi-Agent Sycophancy Is a Pretrained Mid-Layer Vulnerability

Abstract

Multi-agent LLM pipelines route model outputs between instances, where a peer asserting a wrong answer can flip a subject model from correct to incorrect at rates we term *yield*. This is widely attributed to RLHF-induced sycophancy, which would make better post-training the fix. We test this across four model families and find standard alignment is not the cause: pretrained base models exhibit the same substitution pattern as their Instruct variants and usually yield more, and which model family a system is built on explains roughly nine times as much of the variation in yield as whether that model was instruction-tuned. Activation patching on Llama-3.1-8B-Instruct and on its pretrained base localizes the corruption to the same narrow mid-layer window, where attention carries the causal weight and MLP contribution is negligible; patching above it restores 96% of the clean-to-pressured P(correct) gap on Instruct and at least 88% on base. On both, activation-space interventions show that pressure suppresses clean-reasoning features rather than activating a new sycophancy circuit. The attack surface decomposes into two independent factors (channel framing and consensus strength) whose interaction produces a 47.5 percentage-point yield gap at majority consensus, preserved across jury sizes ; yield is highest through tool returns, the channel deployed function-calling pipelines use. A single correctly-arguing dissenter reduces yield by more than 50 percentage points across all framings tested, whereas the strongest prompt-level defense fails on attack variants outside its design surface. Because the vulnerability is pretrained and mid-layer, mitigations should target the mechanism, structured dissent at the pipeline level, rather than prompt-level defenses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.