acceptodds
Under review as a conference paper at ICLR 2027

True Context Corrupts Known Facts in Language Models

Abstract

We study the behaviour of LLMs on structured puzzles: each puzzle consists of a candidate answer, and a list of categories, and the LLM is asked the yes/no question of if the candidate lies in the intersection of these categories. By exploring the case wherein the candidate answer fails membership in a single ‘disqualifier’ category, but lies in the rest, we establish that LLMs are ‘distractible’: even when in isolation the models are able to correctly reject membership in the disqualifier, as we add more correct distractor categories, they progressive become more likely to incorrectly say ‘Yes,’ to the extent that with a few distractors, the majority of puzzles are answered incorrectly. This distraction is also shown to arise in free-form generation. We further establish that ‘fact corruption’ plays a non- trivial role: in certain settings, distraction arises not because the LLMs cannot conjunct category-memberships, but because they begin to endorse membership in the disqualifier. Together, this describes a new failure mode of LLMs on factual problems, and suggests a plausible mode by which hallucinations may occur.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.