Retrofitting LLM-Based Pathway Enrichment Analysis with Dempster-Shafer Theory of Evidence
Abstract
Pathway enrichment analysis (PEA) attaches functional meaning to sets of observed biological features, such as transcriptomic or proteomic measurements, using curated, labeled knowledge bases. Although PEA is a cornerstone of biological discovery, it still relies on relatively simple statistical tests that cannot exploit the flexibility of modern machine learning tools such as large language models (LLMs). Recent attempts to apply black-box LLMs to PEA have had limited success: LLMs struggle to deliver reliable answers over long, heterogeneous inputs, and current remedies (for example, consensus ranking over many partial queries) incur exploding token costs and break the independence assumption required for valid aggregation. Here we present PEACK (Pathway Enrichment Analysis Curated via Knowledge), a post hoc PEA method that corrects black-box LLM responses using common prior knowledge, such as the STRING interaction network (almost any structured prior suffices), and requires no additional LLM calls. Identifying the core failure of partial-query methods as a "Collider Bias", we correct this bias with the Dempster-Shafer theory of evidence, which represents the uncertainty left by a partial query explicitly rather than counting it as evidence against a pathway. We evaluate four frontier open-source LLMs under five PEA query setups on 15 datasets. PEACK significantly outperforms all existing baselines, raising precision-recall AUC for all four LLMs from 0.14-0.38 to above 0.90, which to our knowledge is the best reported PEA performance for those datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.