ScienceDoc: Multidisciplinary Multimodal Evidence-Grounded Question Answering for Scientific Documents
Abstract
Multimodal large language models (MLLMs) are increasingly applied to document understanding and scientific question answering (QA). However, current evaluation remains limited by narrow disciplinary coverage, insufficient assessment of multimodal evidence grounding, and limited evaluation of cross-page and multi-document reasoning. We introduce ScienceDoc, an evidence-grounded multimodal question answering benchmark built from 693 scientific documents, containing 2,200 questions across 8 major disciplines and 61 scientific fields. ScienceDoc includes 4 task types, namely General, Unanswerable, Reasoning, and Multi-Document, together with 21 discipline-specific question types, covering multimodal content from text, equations, tables, and figures. All initial questions are designed by human experts, with expert-annotated answers and manually labeled evidence pages, supporting high-quality hierarchical evaluation of answer correctness and evidence grounding. Experiments on 11 representative MLLMs reveal challenges in multi-document QA and large performance variation on unanswerable questions. The study findings also indicate significant differences across scientific disciplines and gaps between answer accuracy and evidence localization, suggesting that conventional answer accuracy alone is insufficient for evaluating trustworthy scientific document question answering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.