Learning Multiome Representations Across Individuals
Abstract
Paired single-cell multiome assays measure gene expression and chromatin accessibility in the same cell. Contrastive learning on the two measurements has therefore become a standard way to learn cell representations. In population studies, however, such representations are useful only if they transfer across donors. We show that the contrastive loss alone does not ensure this. When negatives are sampled from all donors, the contrastive loss rewards features that identify the donor. When negatives are drawn within donors, each donor's representation can be rotated separately without changing the loss. Removing donor information altogether resolves neither issue without also discarding phenotype. We propose Minuet, which combines within-donor contrastive learning with a prediction task shared across donors and a donor penalty applied within groups of donors that share biological covariates. We show that these objectives determine the representation up to one shift for each group of donors connected by shared covariates. Applied to six paired cohorts, Minuet achieves the best average rank among nine methods in integration, cross-donor and cross-dataset transfer, and phenotype prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.