Misled by the Few: A Mechanistic Analysis of Demonstration Conflicts in In-Context Rule Inference
Abstract
In-context learning enables large language models to perform novel tasks through few-shot demonstrations, including in real deployments where models must infer underlying rules from examples alone. However, demonstrations can naturally contain conflicting examples, making this capability vulnerable. Thus, understanding the dynamics under such conflicts is essential for reliable ICL deployment. In this work, we investigate this question on tasks requiring genuine demonstration reliance. We find that models suffer substantial performance degradation from a single demonstration with a corrupted rule. This systematic misleading behavior motivates our investigation of how models process conflicting evidence internally. Using linear probes and logit lens analysis, we discover that under corruption, models encode both correct and incorrect rules in intermediate layers but develop prediction confidence only in late layers, revealing a two-phase computational structure. We then identify attention heads for each phase underlying the reasoning failures: Vulnerability Heads in early-to-middle layers exhibit positional attention bias with high sensitivity to corruption, while Susceptible Heads in late layers significantly reduce support for correct predictions when exposed to the corrupted evidence. Targeted ablation validates our findings, with masking a small number of identified heads improving accuracy by up to 11.12% relative to the corrupted baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.