acceptodds
Under review as a conference paper at ICLR 2027

Scenario Collapse: Measuring Scenario Local Evidence Binding in Long Context Reasoning

Abstract

Large language models can retrieve the right evidence from a long context yet apply its value to the wrong case or document version. We call this failure scenario local binding. We introduce ScenarioCollapseBench, a paired scenario benchmark with 10,132 scenario pairs and deterministic evidence, state, and decision traces, together with found-but-not-bound (FBNB), which measures neighboring scenario value substitution after successful evidence localization. Across ten models, substantial binding errors remain even after the correct evidence is found. The same failure appears in 200 naturally occurring historical versions of UK legislation and paired adaptations of RULER and BABILong. A matched Qwen3-14B thinking ablation shows that additional reasoning reduces, but does not eliminate, the problem. Oracle experiments locate the dominant residual error in value/state construction rather than evidence access. Finally, making scenario local state explicit through scenario specific identifiers, state tables, and verification substantially reduces binding errors and collapse. These results identify scenario-local binding as a distinct reasoning problem and provide a practical design principle for systems that reason across similar cases, versions, or branches.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.