acceptodds
Under review as a conference paper at ICLR 2027

Beyond Gene Counts: Self-Supervised Cell Representations from Individual Molecules

Abstract

Image-based spatial transcriptomics records the position of every mRNA molecule inside each cell, yet representation learning for these data still reduces each cell to a vector of gene counts, discarding the subcellular organization of these molecules. Where transcripts sit within a cell, near the nucleus, along the membrane or in protrusions, varies between cells in ways a count vector cannot record. We formulate subcellular representation learning as self-supervised learning over point clouds of molecules, each labeled with one of thousands of genes, observed in an arbitrary orientation and with incomplete detection. We present Molm, the first self-supervised model for this setting, which represents a cell as a hierarchy of molecules, the localization programs they form, and the cell as a whole. At every layer, a transformer groups molecules into programs, and the geometry of each program is computed exactly, respecting rotations and reflections by construction. We also introduce a benchmark with known subcellular ground truth, built on real cells whose molecule positions are resampled from known localization programs, and use it, with three real datasets from different imaging platforms, to evaluate nine existing methods. Molm recovers a cell subtype defined only by molecular arrangement at 0.780 balanced accuracy against 0.426 for the best existing method, and localization patterns at 0.448 against 0.259. On real data it leads all label-free methods in cell-type prediction while additionally capturing cell shape, which count-based and foundation-model embeddings do not encode.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.