BioTraceBench: Evaluating Computational Execution and Evidence-Based Reasoning in Biomedical Research
Abstract
Assessing AI for biomedical research requires distinguishing the ability to execute computational analyses from the ability to interpret experimental evidence and synthesize scientific conclusions. We introduce BioTraceBench, a benchmark that evaluates these capabilities through three independently administered tasks: computational execution, local evidence interpretation, and study-level conclusion synthesis. The benchmark comprises 60 curated study entries and 1,061 local conclusion nodes spanning three biomedical research domains. Twenty-four domain researchers constructed structured nodes through schema-guided annotation and expert review, linking local conclusions to experimental designs, results, essential context, and source evidence. Task 1 requires models to execute bioinformatics analyses and produce verifiable computational artifacts. Task 2 provides 384 expert-reviewed, minimally sufficient evidence packages combining experimental figures or tables with essential textual context, from which models infer local conclusions. Task 3 requires models to integrate local conclusion nodes and study background into self-contained study-level conclusions. During construction, three senior doctoral researchers conducted a blinded review of all 60 entries to assess whether the supplied nodes support the target inference using domain knowledge. Each task separates model-visible inputs from evaluation references and uses task-specific scoring. Expert ratings additionally assess correspondence between Task 2 automated judgments and human assessments. Evaluation reveals remaining limitations in computational execution, local evidence interpretation, and study-level conclusion synthesis. The tiered Task 3 evaluation distinguishes recovery of core conclusions, preservation of appropriate inference scope, and citation of sufficient supporting nodes, showing that recovering a core conclusion does not necessarily satisfy the accompanying scope and support requirements. By relating task performance to explicit input conditions and scoring requirements, BioTraceBench provides a traceable basis for diagnosing specific capability gaps in biomedical AI and guiding model improvement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.