acceptodds
Under review as a conference paper at ICLR 2027

Crossed-Witness Grounding for Evidence-State-Guided Visual Document QA

Abstract

Partial evidence makes document question answering a question-relative decision: the same incomplete page can retain a direct witness for one question while losing it for another. Marginal scores on separately altered packets do not establish whether the same incomplete packet elicits opposite question-conditioned decisions. We introduce Crossed-Witness Grounding (CWG), a source-bound OCR intervention, reader-adaptation, and reciprocal-audit protocol. CWG pairs two questions with distinct witnesses on one page and removes either witness in turn. Each altered packet requires one exact answer with a citation to a visible block and one explicit refusal; the two deletions reverse these decisions. At deployment the adapted reader sees one question and its visible packet. On 77 source-disjoint held-out pairs, a frozen Qwen3-VL-8B reader passes 3 (3.9%) four-decision tests, paired adaptation passes 26, and the full training sequence passes 33 (42.9%). On 308 original questions, paired adaptation improves joint intact-answer and own-witness-removal behavior by 32.1 percentage points. A 2B reader shows the paired-training gain on the same protocol. Answer-preserving deletion exposure further improves the 8B four-decision score over matched intact continuation. CWG provides a strict reader-level test of question-conditioned evidence use on incomplete known-page packets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.