acceptodds
Under review as a conference paper at ICLR 2027

Reward-Guided Latent Diffusion for Scalable Feature Selection

Abstract

Feature selection (FS) is a cornerstone of high-dimensional data analysis, aimed at identifying informative feature subsets that enhance model generalizability and interpretability. Despite the emergence of latent-space generative feature-selection frameworks, two challenges remain: candidate generation can rely on direct optimization of selected latent points, and dense interaction modeling becomes prohibitively expensive as dimensionality grows. We propose Reward-Guided Feature Selection (RGFS), which reformulates FS as a reward-guided latent denoising process. By integrating a Latent Diffusion Model (LDM) with a Variational Autoencoder (VAE) and a Reward Evaluator, RGFS models the empirical distribution of latent embeddings from collected reward-labeled masks. During inference, evaluator gradients steer independently initialized reverse trajectories toward regions of higher predicted utility. To further scale interaction modeling to ultra-high-dimensional data, we introduce RGFS+, which incorporates Graph-Induced Sparse Attention. RGFS+ restricts each feature to a correlation-defined neighborhood, reducing sparse attention computation and attention-weight storage from to for fixed neighborhood size . Across 14 standard benchmarks, RGFS or RGFS+ achieves the highest or tied-highest reported score. On two additional datasets with approximately 60,000 features, RGFS+ remains runnable when the evaluated dense alternatives encounter out-of-memory errors. These results support reward-guided diffusion for feature-mask generation and graph-induced sparse attention for high-dimensional feasibility under the evaluated settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.