Fair Subset Selection across Categories and Views via Adaptive NEPv
Abstract
Modern machine learning increasingly relies on large datasets, but limited computational budgets can make training on all available examples impractical. In many cases, selecting a small training subset requires balancing category representation against feature diversity. This is especially difficult when categories overlap and one item contributes to several categories under a shared budget. We construct positive-definite category kernels by softening membership weights on a common feature geometry, then maximize a soft minimum of normalized log-determinants. Endpoint identities separate membership-count effects from diversity and distinguish zero-ridge and ridge-dominant regimes. A Stiefel relaxation yields an adaptive nonlinear eigenvalue formulation and a self-consistent-field solver. Its symmetric lifting has rank at most twice the subset size, an explicit spectrum, a uniform eigengap, and exact small-matrix compression; local convergence remains conditional on operator sensitivity. At a budget of 500 examples, trademark experiments report 81.91% mAP on OpenLogo retrieval and 53.67% category recall on USPTO description generation, compared with 76.15% and 51.36% under uniform sampling. The method does not lead on every metric. We distinguish these empirical comparisons from hard category quotas, global optimization guarantees, and downstream coreset guarantees.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.