acceptodds
Under review as a conference paper at ICLR 2027

Snapshot Illusion: Cursor Exhaustion Does Not License Global Claims

Abstract

Tool-using language-model agents often treat cursor exhaustion as evidence for an exact global claim. Yet complete pagination certifies a traversal, not that its pages describe one database state. We formalize this Snapshot Illusion as an identifiability failure: if two executions induce the same trace and contract but different target answers, no trace-only rule can soundly choose an exact answer. The result covers counts, sets, none/all, uniqueness, sums, and extrema. Across 7,220 versioned SnapshotBench episodes, five open-weight families make unsupported exact claims in all 719/719 valid controlled dynamic responses; matched eligible controls and calculator normalization rule out traversal and arithmetic as the main cause. The rate is 238/300 on public NYC 311 histories, 67/420 for seven advanced APIs on frozen traces, and 86/120 under ordinary six-model interaction. SnapGuard, a certificate-checking release monitor, has 0/880 observed unsafe structured releases and 755/880 safe resolutions. On the same saved natural responses, blinded review finds that a heuristic adapter leaves 18/120 unsafe releases and yields 90/120 valid qualifications. Target-bound snapshots permit exact release in 106/108 eligible runs, but only 58/120 scheduled runs are correct. Capability narrows the empirical gap; coherent evidence closes the logical one but does not ensure correct computation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.