PaperMolBench: Benchmarking and Understanding Molecular Structure Recognition in Scientific Figures
Abstract
Scientific figures in chemistry and biology contain complex molecular structures and figure logic. Existing benchmarks mainly focus on single-molecule images and therefore do not measure whether visual language models (VLMs) can understand real scientific figures for research in these domains. We introduce PaperMolBench, a human-annotated benchmark for extracting target molecular structures and reactions from full scientific figures, comprising 355 test cases. We evaluate representative frameworks, including Tool-Cascaded Method, Agentic Pipeline, and End-to-End Generation, together with several widely used large models. On Molecule Extraction, the best configuration among evaluated methods reaches only 50.0% exact-match (EM) accuracy. To understand the gap, we build a capability diagnostic framework for the simple yet competitive end-to-end generation approach, covering localization, identification, and transcription. Using InternVL3.5-8B, we observe low accuracy even on synthetic molecular images, with further drops on crops from scientific figures, revealing a clear distribution gap. Further analysis indicates that effective molecular structure extraction requires addressing both limited data scale and distribution mismatch, and that linear text representations such as SMILES do not naturally capture molecular topology. These findings highlight key directions for future VLMs in chemistry.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.