Simple Siamese Networks on Foundation Model Embeddings Outperform Complex Models on Unseen Synthetic Lethality Prediction
Abstract
Synthetic lethality (SL) is a biological phenomenon whereby co-deletion of two genes results in cell death, whereas deletion of either has minimal effect, and is widely considered a promising route for cancer therapeutic target discovery Over the past decade, various machine learning methods have been proposed to predict synthetic lethal interactions from gene-level features. However, recent advances in AI for biology have produced a rich set of pretrained embeddings, each capturing distinct multi-modal representations defining genes and proteins, which have as-yet not been integrated into SL prediction models. Here, we propose a lightweight deep learning framework that uses multiple pretrained embeddings for the prediction of SL. On the largest database of SL interactions, we show such lightweight models achieve comparable performance to existing methods despite having orders of magnitude fewer trainable parameters, while the advantage of graph-based methods on previously seen genes disappears under strict gene holdout, likely due to existing SL methods leveraging the SL graph itself. We conduct a comprehensive ablation analysis, finding embeddings of curated knowledge generalize best, and several individual embeddings outperform naive multi-modal fusion. We further find integrating cellular context improves generalization capacity. Finally, we evaluate SL prediction on small focussed screens unseen during training, to our knowledge first evaluation to test whether such models could help actively guide future SL screens. We find poor zero-shot performance across all models, and evaluate fine-tuning strategies possible with lightweight models to improve predictive performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.