acceptodds
Under review as a conference paper at ICLR 2027

Hi-Cell: A Cell-Level Foundation Model for Single-Cell Hi-C Data

Abstract

Single-cell Hi-C (scHi-C) technology provides a direct view of three-dimensional genome organization in individual cells, but its high dimensionality, extreme sparsity, and multiscale contact patterns create an urgent need for a sophisticated and versatile computational framework. Here we introduce Hi-Cell, a cell-level foundation model pretrained on Human-Hi-C-Corpus, our curated large-scale collection of 324,102 single cells from 17 human tissues and 357 bulk Hi-C experiments. Hi-Cell represents each cell as a hierarchical "cell sentence” composed of cell-, chromosome-, and multiscale patch-level tokens, and uses global bidirectional self-attention to integrate local contact patterns with chromosome-wide and whole-cell organization. A three-stage progressive pretraining strategy combining single-cell pretraining, bulk-derived pseudo-single-cell supervision, and cross-depth refinement enables Hi-Cell to learn transferable cellular representations and structural priors. Across six unseen scHi-C datasets, zero-shot embeddings are competitive with dedicated methods, while label-free fine-tuning achieves the best mean clustering performance among the evaluated baselines. Frozen Hi-Cell representations also support accurate cell-type annotation, achieving the highest mean macro-F1 among baselines. Beyond cell representation, Hi-Cell performs zero-shot contact imputation across sequencing depths, leading all evaluated methods in held-out and leave-one-out structural correlation on all six datasets and preserving independently validated A/B compartment profiles. Together, Hi-Cell provides a unified and transferable framework for analyzing cellular identity and reconstructing multiscale 3D genome organization from sparse scHi-C data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.