acceptodds
Under review as a conference paper at ICLR 2027

Cross-Modality Representation Interpretation for Single-Cell Language Models

Abstract

Single-cell foundation models continue to improve on standard benchmarks, but it remains unclear whether their internal representations encode biological knowledge or merely capture subtle expression patterns correlated with benchmark labels. We investigate this question in Cell-to-Sentence-Scale models, which represent cells as gene-name sentences and thus provide a natural setting for examining whether language-style training helps models capture biological knowledge. Using two human peripheral blood mononuclear cell datasets, we train sparse autoencoders (SAEs) on gene-level residual-stream activations from several intermediate layers, and evaluate the learned feature dictionaries along three axes: representation quality, downstream utility, and biological interpretability. These learned features are broadly utilized, predominantly selective, and nonredundant, with these properties persisting under distributional shifts in the input data. Replacing gene embeddings with their SAE reconstructions largely preserves downstream task performance, suggesting that the information discarded by the SAEs has limited relevance to the tasks evaluated. Gene Ontology enrichment analysis further reveals that feature semantics sharpen with layer depth, progressing from generic immune to lineage-specific programs. Steering with interpretable biological concepts shifts cell-type predictions in the intended directions, providing further evidence that the foundation model learns biologically meaningful representations. This control varies across layers, while SAE-reconstructed representations preserve the model’s responsiveness to feature-level perturbations. Together, these findings provide a practical framework for interpreting cross-modal single-cell foundation models and diagnosing their learned representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.