Evidence Over Memory: Do Long-Context LLMs Reason from What They Read?
Abstract
Long-context models can read an entire scientific paper, but do their conclusions actually depend on what they read? We introduce MainLogic, a benchmark that tests selective responsiveness: a model should change its judgment when decisive evidence changes, withhold judgment when evidence becomes insufficient, and remain stable when irrelevant content changes. Starting from human-verified scientific citations, we construct matched families that systematically vary the claim and its evidential basis while preserving the underlying scientific problem. Across ten frontier models, performance on natural claims substantially overstates evidence-grounded reasoning: models reaching over 90% accuracy on the original claims correctly solve at most 32.4% of complete evidence-intervention families. Moreover, when supplied evidence conflicts with a model’s independently measured prior, accuracy drops by 6–51 points, and errors systematically drift toward that prior even when the evidence is insufficient. MainLogic reveals a gap between reading evidence and letting evidence determine the answer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.