When Evidence Disagrees: Shapley-Supervised Evidence Selection for Perturbation Response Prediction
Abstract
Predicting gene expression responses to perturbations is essential for elucidating gene function and guiding experimental design. Recent methods draw on paths retrieved from biological knowledge graphs (KGs) and measurements from related experiments to predict unobserved responses. However, signed KG paths can imply opposing effects, while experimental measurements can report different responses across conditions. The usefulness of each path or measurement can also depend on the perturbation, target gene, cell line, other evidence, and the downstream predictor that uses it. To address this challenge, we introduce PertSelect, a predictor-specific evidence selector that learns which retrieved KG paths and measurements to provide to a given predictor. During training, we estimate predictor-specific Shapley values by averaging how adding each evidence candidate to different subsets changes the predictor's probability of the correct response. Its three variants, PertSelect-Joint, PertSelect-Independent and PertSelect-Kind, use ModernBERT to learn contribution scores and retain evidence with positive scores at inference. We evaluate PertSelect with two predictors on CRISPR interference (CRISPRi) Perturb-seq data from four cell lines: Qwen3-8B, which reads KG paths from OmniPath and INDRA/CoGEx together with related experimental measurements, and a voting predictor that aggregates observed responses of the queried target gene in related perturbation experiments. Across 35,170 held-out queries in four evaluation settings, PertSelect-Joint improves macro-F1 over providing all retrieved evidence by 20.3–31.5 percentage points for Qwen and 8.6–14.4 points for voting. It also achieves higher macro-F1 than embedding-similarity ranking, LLM-based pruning and a learned filter in the style of FILCO in every setting for both predictors. For Qwen, most of its gain comes from deciding how many KG paths and measurements to provide, which removes a strong bias toward predicting down-regulation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.