CREST: Contrastive Metric Learning for Perturbation Modeling and Retrieval
Abstract
In target and drug discovery, the vast combinatorial space of genetic and chemical perturbations makes exhaustive screening intractable, positioning candidate retrieval – the in silico prioritization of promising conditions for physical validation – as a central practical application. Yet current models frequently fail to generalize across unseen cellular contexts, suffering from mode collapse, a behaviour which is encouraged by pointwise reconstruction losses that lack geometric constraints to separate distinct perturbation signatures. To address this, we present CREST (Contrastive Retrieval of Expression State Transitions), a metric learning framework that co-embeds context-conditioned perturbation queries and measured transcriptional responses in a shared space via contrastive alignment. On LINCS L1000 profiles, CREST substantially outperforms all evaluated baselines on candidate retrieval under cell line holdout, and as context diversity scales up to 86 training cell lines, CREST exhibits near-linear gains in retrieval, while capacity-matched pointwise models remain nearly flat. CREST’s discriminative space flexibly extends to transcriptional response prediction: retrieving nearest-neighbor reference profiles for unseen queries yields response estimates that outperform dedicated models. Moreover, simply swapping the reference profiles at inference enables the frozen, bulk-trained CREST to transfer zero-shot to single-cell Perturb-seq, outperforming specialized single-cell models across unseen conditions. Overall, CREST establishes contrastive metric learning as a scalable foundation that unifies candidate prioritization and perturbation response prediction within a single geometric framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.