True Context Corrupts Known Facts in Language Models
Abstract
We study the behaviour of LLMs on structured puzzles: each puzzle consists of a candidate answer, and a list of categories, and the LLM is asked the yes/no question of if the candidate lies in the intersection of these categories. By exploring the case wherein the candidate answer fails membership in a single ‘disqualifier’ category, but lies in the rest, we establish that LLMs are ‘distractible’: even when in isolation the models are able to correctly reject membership in the disqualifier, as we add more correct distractor categories, they progressive become more likely to incorrectly say ‘Yes,’ to the extent that with a few distractors, the majority of puzzles are answered incorrectly. This distraction is also shown to arise in free-form generation. We further establish that ‘fact corruption’ plays a non- trivial role: in certain settings, distraction arises not because the LLMs cannot conjunct category-memberships, but because they begin to endorse membership in the disqualifier. Together, this describes a new failure mode of LLMs on factual problems, and suggests a plausible mode by which hallucinations may occur.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.