acceptodds
Under review as a conference paper at ICLR 2027

AN APPROACH TO QUANTIFYING THE IMPACT OF ARTIFICIAL INTELLIGENCE ON THE SCIENCES: PERSPECTIVES FROM BIOMEDICINE

Abstract

Machine learning (ML) is now ubiquitous in the biological sciences, and unless implemented carefully it is prone to data leakage and related modes of failure. Estimating how often this happens has so far required manual, domain-specific systematic reviews wynants2020prediction,roberts2021common,whalen2022navigating, which do not scale beyond a few hundred papers per study. Large language models (LLMs), combined with established taxonomies of ML failure modes kapoor2023leakage,lones2024avoid, make a literature-scale estimate tractable khraisha2024can. We applied a lightweight LLM pipeline, a non-ML comparison arm, and a sandboxed code-execution audit to biomedical papers retrieved from OpenAlex — of them classified as using ML — and followed every code-availability claim in the cohort to its source. Classification was checked against human review of a blind, stratified -paper validation set, on which the pipeline was accurate ( CI –), or (–) once review articles are recoded under the pre-specified rule. Within the ML corpus, of papers reported performance metrics inadequate to the claim they were offered in support of, and showed evidence of illegitimate features. Of papers claiming available code, () had a link that still resolved, released code matching the paper's own description of its methods, and exactly one could be built and executed in an isolated sandbox — where it did reproduce its reported result. ML papers nonetheless accrued more citations than non-ML papers after adjustment for publication year and subfield (incidence rate ratio , CI –) and appeared in venues with roughly the median impact factor, and papers whose availability claims failed the audit suffered no citation penalty. This is, to our knowledge, the first literature-scale estimate of ML implementation quality in biomedicine, and it suggests that ML is subject to hasty over-adoption there, with reproducibility and code availability routinely overstated.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.