acceptodds
Under review as a conference paper at ICLR 2027

LEXA: Learning a lexicon of gene-programme states from single-cell transcriptomes

Abstract

Single-cell transcriptomic modeling faces a persistent trade-off between preserving complex cellular variation and providing biologically interpretable representations. Continuous latent models capture rich cellular information but are difficult to interpret at the level of individual dimensions, whereas factorization- and topic-based gene-programme methods provide interpretable components but retain less fine-grained variation. We propose LEXA, which couples programme-factorized vector quantization with scaffold-guided gene responsibility and gene-resolved additive decoding to organize cellular variation into programme-specific axes, where each cell selects one of multiple discrete states along each axis, forming a compositional lexicon with explicit gene-resolved molecular effects. We pretrained LEXA on CELLxGENE and evaluated on multiple held-out datasets against representative continuous, discrete, and gene-programme baselines, and found that it preserves substantially more cellular information than conventional gene-programme methods while yielding state-level molecular interpretations that are biologically coherent and reproducible. In an external COVID-19 cohort, both the atlas-pretrained and cohort-specific LEXA settings revealed severity-associated programme variation across immune lineages beyond conventional cell-subtype descriptions. These results demonstrate LEXA as a framework for reconciling information-rich representation with explicit molecular interpretation, turning cellular variation into enumerable and biologically testable molecular states.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.