Program State Distribution Regularization for Learning Across Single Cell Technologies: Benefits and Limitations
Abstract
Single-cell integration aims to enable joint analysis of datasets by reconciling technical distribution shifts while preserving biological variation. Most methods rely on correspondences or invariances across datasets that can be ambiguous, unstable, or absent when biological support set differs across datasets. Methods based on gene programs offer a more structured unit of alignment, but matching program identity alone does not ensure comparable program states or their combinations. We posit the joint distribution of shared program states as the cross technology invariant, and explicitly decompose it into program specific state distributions and interprogram dependence. We implement this decomposition using a shared program bottleneck with inference and observation modules specific to each technology, aligning program marginals through quantiles and interprogram dependence in rank space. Controlled simulations with five independent generators and five optimization repeats show higher correlation between matched factors than the same model without alignment at all four tested levels of observation distortion. On a development cohort, adding dependence alignment improves mean clustering and transfer scores over marginal alignment across five paired fits; for donors excluded during training, prediction of protein from RNA raises median protein from 0.187 without alignment to 0.250 with the combined objective. These results show that aligning program state distributions can improve recovery of shared factors and prediction across technologies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.