acceptodds
Under review as a conference paper at ICLR 2027

VisProve: Learning Typed Visual Provenance for Scientific Document Question Answering

Abstract

Visual citations show where an answer is supported, but not how the cited content contributes to it. In scientific document question answering, a model may read a value directly, compress several observations, or combine evidence to infer a conclusion. We study typed visual provenance, where a system generates an answer together with source regions and explicit support relations. We introduce **ReViE**, a dataset of 3,941 questions over complete scientific documents, with page, section, and region annotations labeled as *Quotation*, *Compression*, or *Inference*. We further propose **VisProve**, which augments joint answer and provenance generation with differentiable Spatial Evidence Reading and Evidence–Answer Relation Feedback. VisProve conditions next-token prediction on localized visual evidence and its estimated support relation through a gated norm-bounded residual. On the document-disjoint ReViE test set, VisProve improves Typed-F1 from 3.52 to 5.48 and Section-F1 from 45.99 to 53.24 over joint supervised fine-tuning with the same backbone, while Page-F1 changes only marginally. These results suggest that coupling visual grounding with support relation modeling mainly improves fine-grained provenance, while the low absolute Typed-F1 reveals a substantial gap between answer generation and precise visual evidence attribution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.