acceptodds
Under review as a conference paper at ICLR 2027

Evidence Over Memory: Do Long-Context LLMs Reason from What They Read?

Abstract

Long-context models can read an entire scientific paper, but do their conclusions actually depend on what they read? We introduce MainLogic, a benchmark that tests selective responsiveness: a model should change its judgment when decisive evidence changes, withhold judgment when evidence becomes insufficient, and remain stable when irrelevant content changes. Starting from human-verified scientific citations, we construct matched families that systematically vary the claim and its evidential basis while preserving the underlying scientific problem. Across ten frontier models, performance on natural claims substantially overstates evidence-grounded reasoning: models reaching over 90% accuracy on the original claims correctly solve at most 32.4% of complete evidence-intervention families. Moreover, when supplied evidence conflicts with a model’s independently measured prior, accuracy drops by 6–51 points, and errors systematically drift toward that prior even when the evidence is insufficient. MainLogic reveals a gap between reading evidence and letting evidence determine the answer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.