acceptodds
Under review as a conference paper at ICLR 2027

Keep the Question: Task-Preserving Memory for Visual PDF Reading

Abstract

Visual document question answering often decomposes a question into local reading steps. This decomposition can lose part of the original task: a correct intermediate observation may replace a required comparison, count, or sum. We study this failure as task scope loss and introduce task-preserving memory to connect search-time observations with subsequent reading. An initial record links factual requirements to their prerequisites, candidate values, and source pages. The same record guides follow-up retrieval and is supplied to both the rereading planner and final answer generator, allowing local inspection without discarding previously observed facts. The method uses frozen models and requires no additional pages or reader calls. On 212 questions from 106 held-out MMLongBench-Doc documents, retaining the record improves focused rereading from 44.8% to 49.5% local semantic accuracy, a 4.72-percentage-point gain with a 95% document-bootstrap interval of [0.47, 8.96]. A factorial comparison of memory retention and reading policy, record-content analysis, and a LongDocURL transfer evaluation examine the contribution of retained information. The results show that preserving search-time facts can improve task completion under a fixed page and reader-call budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.