PubTables-QA: A Multi-Page Table Visual Question Answering for Cross-Page Table Reasoning
Abstract
Tables are a central medium for communicating quantitative evidence in scientific documents, yet their compact relational form makes robust interpretation depend on layout and context. However, existing multi-page and multi-table benchmarks may still allow stochastic shortcut search, where models locate local answer-bearing regions without systematically navigating the underlying table structure. To address this gap, we introduce PubTables-QA, a multi-page TableVQA benchmark dataset carefully designed to audit cross-page visual reasoning in authentic academic documents. Built upon the extensive structural table annotations of PubTables- v2, PubTables-QA introduces 1,871 table-based QA pairs, where each question is designed to be unanswerable from a single page alone. This design compels models to reconstruct cross-page table continuity across headers, rows, columns, and cells, or to integrate evidence distributed across multiple tables. PubTables-QA further provides a three-level fine-grained QA taxonomy: Document, Table, and Cell/column levels respectively, ranging from relevant evidence localization and table-level structural understanding, to precise cell/column-level value extraction. Experiments on recent MLLMs, including GPT-4o, Gemini-2.5-Pro, and Gemma-4 reveal that even the strongest models achieve 45.9% overall accuracy. Further, our evidence-aware evaluation confirms that, unlike existing multi-page benchmarks susceptible to stochastic shortcuts, PubTables-QA requires genuine cross-page evidence integration, establishing it as a substantial and unsolved challenge. Code and Dataset are available at https://anonymous.4open.science/r/PubTables-QA-584D.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.