What Do Transcriptome Models Learn? From Expression Probing to Biological Prediction and RNA Question Answering
Abstract
Transcriptome foundation models are increasingly used for biological prediction and connected to language models to answer questions about transcriptomic data. However, differences in both pretraining methods and training data make it difficult to determine which approach works best for a given use case. To address this, we train six single-cell and four bulk models from scratch on one shared corpus per modality, following literature-inspired pretraining recipes. Our evaluation addresses three questions. *How does pretraining affect expression probing?* The recipe with the highest score after pretraining is not always the one that improves the most, and for some recipes linear probes already achieve high scores on representations from untrained models. *How do the recipes perform on biological tasks?* Pretraining improves classification over untrained models for every recipe, but its effect on perturbation prediction is inconsistent. Moreover, recipes that score higher on expression probes do not necessarily perform better on biological tasks. Early in pretraining, LogME transferability scores already correlate positively with recipe performance on all five biological tasks, without training any downstream predictor. *Do the differences observed in probing and biological prediction extend to RNA question answering?* We connect the frozen pretrained models to Qwen and evaluate the probability assigned to reference answers for released CellWhisperer questions. Among single-cell models, cell-type classification rankings agree most closely with reference-answer likelihood rankings on cell-identity questions, whereas expression-value probe rankings show no such correspondence. Cell-type LogME rankings also agree between RNA and Qwen representations with an untrained adapter, suggesting that, as in vision–language models, encoder strengths carry over into the language model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.