acceptodds
Under review as a conference paper at ICLR 2027

Identifying Bias Directions from Population Residuals in Reference-shifted Deconvolution

Abstract

We study reference-based linear deconvolution problems, which estimate component fractions from aggregate observations using a potentially shifted component profile. Our study is motivated by gene expression deconvolution. In biomedicine, bulk RNA-seq measures aggregate gene expression in a tissue; to infer cell-type composition, one deconvolutes it by referring to a cell-type gene expression profile, which is yet measured from separate tissues and by a different technology (scRNA-seq). Consequently, the reference profile can significantly differ from the true profile that generated the aggregates. Previous methods either neglect this shift or address it only in preprocessing; they then mostly use least squares to perform deconvolution for each tissue individually. Given that a cohort of similar bulk RNA-seq measurements is usually available, we propose to use post-fit residuals together to correct individual deconvolution bias. Our algorithm is developed from rigorous analysis. We first show that under a shared shift in reference, the population residuals alone cannot identify the fraction biases, motivating us to build a structural bridge that connects markered residuals to the population fraction error. Then, based on the latter, we study the conditions under which individual fraction bias can be safely corrected. The efficacy of our correction algorithm is proven under the PAC framework. Experimentally, we first conduct synthetic studies that demonstrate the residual geometry, bias identifiability, and PAC learning effectiveness. Second, we test an instantiation of our algorithm named PopLS on eight cancer pseudo-bulk datasets and two real-bulk cohorts, which further verify the benefits of population-based deconvolution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.