From gene programs to biological language: rethinking single-cell foundation models as modality translators
Abstract
Single-cell foundation models effectively capture gene expression patterns and cellular heterogeneity to support downstream analysis, but struggle to incorporate broader knowledge. Large language models possess such knowledge, yet cannot directly interpret single-cell expression data. Connecting these complementary capabilities often requires specialized model designs or modifications to existing frameworks. We introduce CellDialect, a framework that repurposes pretrained single-cell foundation models as modality translators. CellDialect converts expression profiles into cell-conditioned continuous tokens compatible with a frozen language model, without pretraining a new language backbone or changing its architecture. Joint numerical and language supervision supports cell representation learning, question answering, descriptive text generation, and expression generation conditioned on cellular metadata. Benchmark evaluations demonstrate improved representation quality over the evaluated baselines and competitive zero-shot cell task question answering. CellDialect connects existing single-cell and language foundation models within a shared framework. It offers a path to supporting diverse single-cell tasks and improving existing models with modest additional pretraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.