scCGNet: Jointly Optimizing Copula Mixtures and Graph Autoencoders for Single-Cell Clustering
Abstract
Clustering is central to single-cell RNA sequencing (scRNA-seq) analysis but remains difficult: the data are high-dimensional, sparse, and correlated. Recent deep methods increasingly model this correlation to learn better representations- for example, through gene-gene structure or masked reconstruction. The clustering step, however, does not. After embedding, cells are grouped by K-means, deep embedded clustering, or contrastive objectives that rely on distances to cluster centroids. These objectives treat the latent dimensions as independent within a cluster, and cannot capture the non-Gaussian, asymmetric, and tail dependence that transcriptomic data exhibit. We propose scCGNet, which models dependence at the clustering stage. A variational graph autoencoder encodes a cell-cell k-nearest-neighbor graph into topology-aware latent representations. A Gaussian mixture copula model then clusters cells by explicitly modeling the dependence structure within each cluster. A zero-inflated negative binomial layer accounts for sparsity and over-dispersion. The three components share one objective and are trained jointly. On fifteen real scRNA-seq datasets, scCGNet attains the highest and most consistent accuracy, and ablations show that the copula-based cluster model drives the gain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.