acceptodds
Under review as a conference paper at ICLR 2027

Not Proven Is Not Proven-Not: Language Models Read the “Not” in a Question as Negation as Failure

Abstract

Language models are routinely asked about texts that leave some facts open, and then the right answer is unknown. We find that a “not” turns such gaps into confident answers: when a text mentions a property P but never settles whether x has it, models judge “x is P” unknown yet “x is not P” true. The models read the “not” as negation as failure, taking “it is not established that x is P” for “it is established that x is not P”. Both questions look P up in the text, and only the negated one reads a failed lookup as an answer. Editing the text moves the error where this account predicts: removing the property word largely removes it, a rule that merely mentions the word brings it back, and writing “not P” into that rule spreads it to the plain question. Inside the model, the lookup runs through the attention of the question's property word to the text, and while answering true the model still registers, and even says, that the fact is not established. Across fourteen open models the asymmetry tracks this lookup, which is present before instruction tuning. Instructions mostly make models abstain, while thinking replaces the lookup with a proof. What is not proven is read as proven not.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.