BioReasonBench: A Diagnostic Benchmark for Multimodal Reasoning with Structured Visual Evidence in Biology
Abstract
Multimodal Large Language Models (MLLMs) can describe scientific images fluently, but it remains unclear whether they can use biological images as structured evidence across a complete reasoning chain. Existing evaluations are often text-only, recognition-oriented, or summarized by a single accuracy that obscures where reasoning fails. We introduce BioReasonBench, a diagnostic benchmark of biology questions, spanning three question types: single-select, multiple-select, and open-ended. Of these, (\(87.0%\)) are grounded in images of pedigrees, pathways, experimental setups, quantitative curves, and other biological content. Unlike the single-answer choice and single-target open responses common in multimodal benchmarks, \(193/194\) multiple-select items have two or three correct options, and all open-ended items are multi-blank ( blanks in total). Each item is organized along five diagnostic axes and source/year splits, and is evaluated with question-type-aware exact, partial-credit, and LLM-judge protocols. Across MLLMs, mean accuracy falls from \(61.2%\) on single-select to \(28.3%\) on exact multiple-select and \(11.6%\) on fully correct open-ended questions. Relation graphs, tables, and genetics are especially difficult. Partial option sets and partially correct blanks are common, revealing failures of correspondence, state tracking, calibration, and causal completeness rather than uniform factual ignorance. On set0, the strongest model scores \(71.0/100\), matching the lowest of Grade-12 students, who average . We hope BioReasonBench provides actionable insights into fine-grained visual-reasoning mechanics and guides the development of next-generation MLLMs capable of trustworthy scientific inquiry.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.