CrossPerturb: A Cross-Context Benchmark and Evidence-Grounded Framework for LLM-Based Virtual Cell Perturbation Prediction
Abstract
Predicting cellular responses to drug perturbations is fundamental to virtual cell modeling and therapeutic discovery. While large language models (LLMs) show growing promise for biological reasoning, existing benchmarks evaluate them solely within the same cellular context. Consequently, simple voting baselines achieve comparable performance to complex LLM reasoning frameworks, obscuring whether models perform genuine biological reasoning or merely exploit majority labels. This limitation highlights the urgent need for more rigorous evaluation benchmarks that test whether LLMs can perform mechanistic reasoning in virtual cell modeling. To bridge this gap, we introduce CrossPerturb, a comprehensive benchmark for evaluating drug perturbation prediction across diverse cellular contexts. Derived from the LINCS L1000 and Tahoe datasets, CrossPerturb assesses transcriptional outcomes across diverse cell lines, drug dosage, and exposure durations. Alongside the benchmark, we introduce ARCS, an evidence-grounded reasoning framework that **A**nalyzes biological pathways underlying perturbations, **R**etrieves cross-context evidence, **C**ritiques candidate evidence, and **S**ynthesizes the final response score through mechanism-constrained aggregation. Extensive experiments demonstrate that our evidence integration framework surpasses current state-of-the-art methods, whereas existing LLM reasoning frameworks struggle to achieve robust performance, confirming that cross-context biological reasoning remains an open challenge for current LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.