SCOPE: Self-Calibrated Ordered Principal Encoding for Tabular Foundation Models in High-Dimensional Data
Abstract
Tabular Foundation Models (TFMs) remain difficult to apply in High-Dimensional, Low-Sample Size (HDLSS) regimes, where thousands of features can exceed practical input limits and dimensionality reduction often requires dataset-specific design or tuning. We introduce SCOPE (Self-Calibrated Ordered Principal Encoding), a deterministic, tuning-free feature compression framework for constructing compact representations of high-dimensional tabular data. SCOPE builds on a global feature ordering produced by LENS (Layout-Enhanced Neighborhood Sequencing), which combines efficient path construction and refinement with a self-calibrated ordering budget. Given this ordering, SCOPE estimates spectral complexity, analytically determines the output dimension, partitions the feature sequence through transition-aware segmentation, and performs local principal projections to obtain a compact representation. Across high-dimensional benchmarks and frozen TabPFN-style backbones, we evaluate the proposed approach on predictive performance, uncertainty-aware ranking, computational efficiency, and sustainability. Results show that SCOPE achieves a favorable compression, prediction, and compute trade-off by reducing representation size, inference cost, fit time, and memory usage while retaining competitive predictive performance on HDLSS benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.