SciAudit: Benchmarking MLLMs for Scientific Error Detection
Abstract
The rapid growth of scientific output makes it increasingly difficult for peer reviewers to examine every paper in depth. As a result, substantive errors or inconsistencies may be missed, leading to monetary, reputational, or scientific costs. Multimodal large language models (MLLMs) could assist scientific review, but existing error detection benchmarks remain limited in data scale, evaluation granularity, and task openness. We introduce SciAudit, a benchmark for open-ended scientific error detection in full papers. SciAudit comprises two complementary subsets spanning 16 scientific domains: Gold contains 2938 human-verified issues from 293 problematic papers, while Silver contains 5441 controlled errors constructed from 527 high-quality papers through a multi-model Draft–Check–Refine pipeline. Each error is annotated with its location, supporting evidence, reasoning, and type. These annotations support quality-aware metrics that assess error detection and the quality of the intermediate reasoning process. We evaluate 13 MLLMs under thinking and non-thinking settings. Results show that current MLLMs remain limited in open-ended error detection, with even the strongest model achieving less than 50% under our quality-aware F1 evaluation. Additional thinking generally improves performance, but errors requiring composite capabilities or multi-step reasoning remain challenging. Furthermore, we find that models can correctly verify 52.5%–81.3% of previously missed errors when explicitly prompted about them. This suggests that many detection failures stem from limitations in error discovery and reasoning, rather than insufficient knowledge alone. We hope SciAudit can provide a rigorous testbed for evaluating MLLMs and developing more reliable agent systems for scientific auditing in the future.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.