GeneICL: A Tabular Foundation Model for Bulk Transcriptomics
Abstract
Gene expression is widely measured in biomedicine, yet clinical outcome prediction remains challenging due to high dimensionality, strong feature correlations, and limited labeled data. Large self-supervised transcriptomic foundation models often fail to outperform simple supervised baselines. Tabular foundation models offer an alternative through in-context learning, but are typically pretrained on generic synthetic data rather than transcriptomic structure. We ask whether transcriptomics-aware pretraining, rather than merely scale, is the missing ingredient. Towards this end, we introduce **GeneICL**, a 4.2M-parameter tabular foundation model that combines a semi-synthetic pretraining prior built from experimentally measured bulk expression profiles with a parameter-efficient weight-tied recurrent architecture. We further enable right-censored survival prediction through a training-free reduction to regression using Cox partial-likelihood residuals. We evaluate GeneICL on a benchmark of 80 clinical outcome-prediction tasks spanning classification, regression, and survival. Tabular foundation models consistently outperform self-supervised transcriptomic models, while GeneICL achieves the best overall rank among all evaluated foundation models and tuned baselines. GeneICL does so with up to 387 fewer parameters, no gradient updates at inference, and inference within seconds on a consumer laptop CPU.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.