When Do Unseen Data Concentrate at Class-Specific Points? Kernel Data Piling of High-Dimensional Small-Sample-Size Data
Abstract
We consider classification problems in which observations are represented by high-dimensional feature vectors while only a small, fixed number of labeled training samples are available from each class. For this purpose, we work under a high-dimension-low-sample-size (HDLSS) asymptotic regime, where the feature dimension tends to infinity while the training sample sizes remain fixed. We study when independent test observations, after being mapped by a kernel classifier fitted to these labeled samples, concentrate at class-specific locations. We identify conditions under which this class-specific test concentration occurs and characterize when the resulting class locations are distinct for several commonly used kernels. We further show that components with very large within-class variation can prevent test-score concentration throughout the empirical span generated by the training sample. To address this failure, we propose a kernel correction that restores separated class-specific concentration under suitable conditions. Simulation studies provide empirical support for the theoretical results developed throughout the paper.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.