Efficient Distributional Encoding for Multi-cellular RNA-seq Data
Abstract
Single-cell RNA-seq models typically tokenize cells independently as sequences of gene or cell-state tokens, causing sequence length to scale linearly with number of cells genes per cell . This scaling is poorly suited to tasks requiring multi-cellular context, such as patient-level disease diagnosis and spatial microenvironment niche mapping, where samples comprise dozens to thousands of cells. Here, we present a benchmarking study of single-cell tokenization strategies for multi-cellular prediction tasks. Holding a standard transformer backbone fixed, we evaluate gene- and cell-level tokenizers across three multi-cellular benchmarks, tracking predictive performance and token usage as the cellular cohort size varies from 1 to 1,000 cells. To overcome linear context expansion, we introduce module distribution encoding (MDE). MDE groups correlated genes into co-expression modules and summarizes the population-level expression profile of each module as a zero-adjusted distribution token. Total sequence length is therefore determined strictly by the number of gene modules rather than cell or gene count, decoupling context length from cohort size while preserving continuous population-level variance. MDE outperforms existing tokenization schemes on disease prediction and niche prediction, and cuts tokens by orders of magnitude, leading to significantly higher efficiency across all tasks. These results characterize the context-accuracy trade-offs of single-cell tokenization and establish sublinear, population-aware encodings as an effective basis for multi-cellular modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.