acceptodds
Under review as a conference paper at ICLR 2027

MMSciFact: A Benchmark for Multimodal Dependency-Aware Scientific Fact-Checking

Abstract

Multimodal large language models (MLLMs) are increasingly used to answer questions about scientific papers with long-form responses grounded in figures, tables, and text. Yet factuality evaluation typically decomposes these answers into atomic claims and verifies them independently. This independence assumption overlooks how observations serve as premises for downstream interpretations, forming a dependency graph within each answer. An interpretation can be supported in isolation yet depend on a premise contradicted by or unverifiable from the source—a failure we call Invalid Support. We introduce MMSciFact, a benchmark for dependency-aware multimodal scientific fact-checking comprising 257 question-answer pairs and 2,737 sentences from 45 papers. Model answers are preserved without post-hoc corruption and annotated with expert grounding judgments, sentence roles, and dependencies. Across 12 MLLM fact-checkers and single-page, cross-page, and full-paper evidence, models struggle to identify statements contradicted by or unverifiable from the source, as well as interpretations based on invalid premises. Providing sentence roles and dependency graphs improves the detection of Invalid Support, but scores remain low even with this information, indicating that dependency structure alone is insufficient for reliable verification. These findings suggest that dependency awareness complements claim-level verification in supporting reliable scientific fact-checking.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.