acceptodds
Under review as a conference paper at ICLR 2027

Efficient Distributional Encoding for Multi-cellular RNA-seq Data

Abstract

Single-cell RNA-seq models typically tokenize cells independently as sequences of gene or cell-state tokens, causing sequence length to scale linearly with number of cells genes per cell . This scaling is poorly suited to tasks requiring multi-cellular context, such as patient-level disease diagnosis and spatial microenvironment niche mapping, where samples comprise dozens to thousands of cells. Here, we present a benchmarking study of single-cell tokenization strategies for multi-cellular prediction tasks. Holding a standard transformer backbone fixed, we evaluate gene- and cell-level tokenizers across three multi-cellular benchmarks, tracking predictive performance and token usage as the cellular cohort size varies from 1 to 1,000 cells. To overcome linear context expansion, we introduce module distribution encoding (MDE). MDE groups correlated genes into co-expression modules and summarizes the population-level expression profile of each module as a zero-adjusted distribution token. Total sequence length is therefore determined strictly by the number of gene modules rather than cell or gene count, decoupling context length from cohort size while preserving continuous population-level variance. MDE outperforms existing tokenization schemes on disease prediction and niche prediction, and cuts tokens by orders of magnitude, leading to significantly higher efficiency across all tasks. These results characterize the context-accuracy trade-offs of single-cell tokenization and establish sublinear, population-aware encodings as an effective basis for multi-cellular modeling.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.