CuBAS: A New Curvature-Based Adaptive Sampling Method for Supervised Classification
Abstract
The informativeness of a training set is often more important than its size. However, existing sampling strategies largely ignore the intrinsic geometry of labeled data. We introduce CuBAS (curvature-based adaptive sampling), a information-geometric framework for adaptive sample selection in supervised classification. The key insight of the new method is that a labeled dataset induces a statistical manifold whose local curvature provides a principled measure of sample informativeness. Specifically, we model label interactions through a -state Potts Markov random field defined on a -nearest-neighbor graph, and derive a closed-form curvature estimator from the observed Fisher information. Adaptive thresholding of the resulting curvature signal partitions the graph into smooth low-curvature regions – well represented by a small number of prototypes, and high-curvature regions concentrated around class interfaces – where samples are most informative for defining decision boundaries. While jointly sampling from both regimes, CuBAS constructs compact yet highly informative training subsets that preserve global representativeness as well as local discriminative structure. Extensive experiments on more than 60 benchmark datasets spanning tabular, image, and biological domains demonstrate consistent and statistically significant improvements over random and entropy-based sampling across a wide range of labeling budgets. Beyond its empirical effectiveness, CuBAS provides a computationally efficient and theoretically grounded connection between information geometry and adaptive data selection for supervised learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.