Decided While Reading: Why a Language Model Stops Using a Fact Placed Before Its Rule
Abstract
Language models are often asked to apply a rule to a few facts and give an immediate verdict, as guardrails, classifiers and judges do. Such prompts place the rule before, among or after the data. They assume that a model with the whole prompt in view applies the rule to the fact it names, wherever the rule appears. We find that this assumption fails. Across 18 open models, moving one unchanged rule sentence just past the fact it names, or after all the facts, sharply reduces that fact's influence on the answer, and the answer follows the facts read after the rule instead. Yet the models still know the fact. Asked separately, they mostly answer correctly what the rule compares and whether the fact matches. Causal experiments explain the gap. As a model reads each fact after the rule, it forms the rule's outcome at that fact's own position, and its direct answer later reads the outcome mainly from those positions. A fact read before the rule has no outcome there, and the answer, which sees both the rule and the fact, does not make up for it. In short, the verdict is decided while reading. For systems that answer directly, the rule must come before the data it governs, and explaining before answering restores the lost fact.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.