How Should We Describe a Gene? Learning Biological Representations for Perturbation Prediction
Abstract
Predicting responses to unseen genetic perturbations requires representing perturbed genes using prior knowledge. Text provides a natural interface to this knowledge, but existing representations rely on fixed, hand-designed descriptions, leaving unclear what to include or how to organise it. We learn a shared description protocol for unseen genes, separating the benefits of richer knowledge from those of its representation. An LLM iteratively revises the protocol using held-out prediction and trial history while the prediction pipeline remains fixed. We predict transcriptome-wide log-fold changes and introduce *Continuous Directed Overlap*, a rank- and direction-aware metric focused on recovering the strongest effects. Across four Perturb-seq datasets, richer knowledge consistently improves prediction, structured descriptions add further gains, and search generally improves validation performance. In our evaluations, the resulting representations outperform the tested gene embeddings across multiple predictors, retain utility across cellular contexts, and enhance predictions from established single-cell generative models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.