acceptodds
Under review as a conference paper at ICLR 2027

Discovering Feature Subspace through Data Learnability: A Kernel-learning Realization

Abstract

Modern machine learning models can fit high-dimensional data despite irrelevant or redundant features. However, discovering a minimal feature subspace that preserves the underlying predictive structure remains challenging. A key difficulty is that feature utility is typically assessed through a particular model, and is therefore entangled with model-specific inductive biases. Here, we revisit the feature subspace discovery problem from a different perspective: the representation-induced learnability of the observed data. Our key idea is that an informative feature subspace should organize data into suitable structures, which benefit model performance across a range of complexities. To quantify this property, we construct kernel-based estimators and track empirical fitting performance as a function of model complexity, measured by the trace of the normalized Gram matrix. This defines a learning curve, and the Area Under the Learning Curve (AULC) provides a scalar measure of data learnability. AULC provides a principled criterion for comparing feature subspaces. Experiments across synthetic and real-world datasets demonstrate its effectiveness in feature subspace discovery and related scientific variable discovery tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.