acceptodds
Under review as a conference paper at ICLR 2027

CLASP: Correlation-Leveraged Across-sample Separability Prioritization for Unsupervised Biomarker Discovery

Abstract

Biomedical data science operates in a regime of high data dimensionality, low sample sizes, and structured noise. As a result, many applications suffer from a reproducibility crisis, in which biomarkers identified in one study fail to replicate in others. To address this challenge, we introduce Correlation-Leveraged Across-sample Separability Prioritization (CLASP), an unsupervised feature selection framework that identifies reproducible feature subsets by maximizing cross-subsample cluster separability. CLASP frames feature selection as a stochastic optimization problem over inclusion probabilities; it uses a Cross-Entropy Method paired with a correlation-augmented Gumbel sampling scheme to find a joint feature set that maximizes cluster distinctiveness. Evaluated both on simulation environments and seven diverse real-world biomedical datasets, CLASP shows state-of-the-art feature selection stability, cluster quality, and out-of-sample cluster stability, thus preserving the reproducibility of selected biomarkers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.