PIE: Generalizing perturbation effects across unseen perturbations, contexts and datasets
Abstract
Predicting cellular response to perturbations is key to understanding biological mechanisms and selecting therapeutic targets. However, generalizing across cellular contexts, perturbations, and experimental datasets remains challenging due to the difficulty of measuring true perturbation effects and incomplete representations of the system being perturbed. We introduce PIE (Perturbation is Everything), which addresses these challenges by reformulating the learning task around population-level perturbation effects and incorporating auxiliary inputs that better characterize the biological system. PIE directly predicts differentially expressed genes and changes in gene expression for each (context, perturbation) pair, using external biological knowledge, baseline gene expression, and observed perturbation responses. On the Replogle-Nadig dataset, PIE achieves state-of-the-art performance on most metrics across all tested generalization settings, including the most challenging setting where both the context and perturbation are unseen during training. PIE achieves 1.2–3.2 times the AUPRC of the strongest baseline in each setting for predicting differentially expressed genes: across unseen contexts, unseen perturbations, and jointly unseen contexts and perturbations. PIE also outperforms existing baselines on zero-shot transfer across experimental datasets. Overall, PIE provides a framework for predicting cellular responses across datasets, contexts, and perturbations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.