acceptodds
Under review as a conference paper at ICLR 2027

How Should We Describe a Gene? Learning Biological Representations for Perturbation Prediction

Abstract

Predicting responses to unseen genetic perturbations requires representing perturbed genes using prior knowledge. Text provides a natural interface to this knowledge, but existing representations rely on fixed, hand-designed descriptions, leaving unclear what to include or how to organise it. We learn a shared description protocol for unseen genes, separating the benefits of richer knowledge from those of its representation. An LLM iteratively revises the protocol using held-out prediction and trial history while the prediction pipeline remains fixed. We predict transcriptome-wide log-fold changes and introduce *Continuous Directed Overlap*, a rank- and direction-aware metric focused on recovering the strongest effects. Across four Perturb-seq datasets, richer knowledge consistently improves prediction, structured descriptions add further gains, and search generally improves validation performance. In our evaluations, the resulting representations outperform the tested gene embeddings across multiple predictors, retain utility across cellular contexts, and enhance predictions from established single-cell generative models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.