acceptodds
Under review as a conference paper at ICLR 2027

Weighted Hybrid PCA: Learning Shared Structure under Batch Effects

Abstract

Principal component analysis (PCA) is widely used to extract low-dimensional structure from high-dimensional data. However, batch effects, such as those due to experimental timing, equipment, or data sources, may dominate the leading principal components and obscure the shared variation of interest. When the corresponding labels are only partially observed, removing these effects without discarding useful observations becomes difficult. We model this setting using a factor model with a shared factor loading space and different intercepts across batches. We then propose weighted hybrid PCA, which combines covariance estimates from labeled and unlabeled observations. Using the observed labels, we select a weight from a safe set that excludes additional components caused by between-population variation. Under suitable conditions, the selected weight asymptotically minimizes both the factor detectability threshold and the factor loading space estimation error over the safe set. The lower threshold allows weaker factors to be detected. Simulations show substantially lower factor loading space estimation error for weights inside the safe set than for those outside it. The selected weight achieves the lowest estimation error, close to the oracle benchmark. On two A549 single-cell RNA perturbation datasets with p53-related expression responses, weighted hybrid PCA more efficiently captures the common within-population variation than the other methods and gives the highest average number and proportion of p53-associated principal components.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.