SILTA: Reference Learning and Target Adaptation for Spatial Transcriptomic Deconvolution
Abstract
Most sequencing-based spatial transcriptomics platforms pool multiple cell-types per capture location, requiring deconvolution to estimate cell-type proportions. Existing methods suffer from spot compositions that are sparse and long-tailed: one type often dominates, yet many methods predict smooth mixtures, and similar gene signatures blur related types. Although neighboring spots share spatial context, spot-wise inference is sensitive to local noise, and these errors accumulate into sample-level composition drift. Platform differences complicate applying composition mappings learned from scRNA-seq references to unlabeled ST targets. To this end, we propose a Two-stage learning model SILTA, Spatial Inference by combining composition Learning and Target Adaptation, which jointly tackles issues above. Stage I learns fine-type boundaries from structured, diffusion-augmented pseudo-spots with the spatial encoder frozen. Stage II refines target proportions from expression and tissue geometry through hierarchical parameter updates, gradual unfreezing, and reference replay. Expression-spatial fusion combines reference composition with target context, while qNB and post-hoc calibration align predictions with measured expression and correct global composition bias. SILTA achieves the best mean for all four primary metrics among evaluated methods on seqFISH, osmFISH and cross-animal Moffitt. We verify both stages' roles through perturbation experiments. Without ground-truth proportions, SILTA demonstrates strong performance in marker-gene concordance across three datasets and in pathology-region classification on HBC, extending its gains to real tissues across assay platforms. The code is at: https://anonymous.4open.science/r/SILTA-SpaDeconv-A38C/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.