Do Large Language Models Treat Rules as Facts?
Abstract
Large language models (LLMs) are increasingly used to reason over policies, procedures, contracts, and audit records. In these settings, we find that models may incorrectly treat rules about what should or should not happen as evidence of what actually happened. We call this failure norm-to-fact projection. We introduce Unstated, a controlled benchmark of 720 instances built from 120 scenarios, pairing unstated-outcome probes with explicit-outcome anchors. Across 28 open-weight models, median accuracy is 84.5% on explicit-outcome anchors but only 6.7% on prohibition probes. Errors are strongly directional: prohibitions are often interpreted as non-occurrence and commands as occurrence. Ten large-scale frontier models perform better overall, but most still achieve below 75% accuracy on the prohibition probes and remain below human performance. Further analyses show that these errors can remain highly confident, persist under family reweighting, and are only inconsistently reduced by prompting. Our results identify norm-to-fact projection as a systematic failure mode and highlight the need to evaluate whether LLMs infer factual outcomes only from factual evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.