Uniform Spectral Bounds under Data-Dependent Sample Selection
Abstract
Concentration inequalities for sample covariance matrices are fundamental tools in high-dimensional probability. However, data-dependent sample selection prevents direct application of classical covariance concentration bounds. In this paper, we study the extreme eigenvalues of covariance matrices formed from arbitrary, possibly data-dependent subsets of random vectors. For i.i.d. Gaussian vectors, we establish simultaneous lower and upper bounds over all subsets satisfying prescribed cardinality constraints. The bounds depend explicitly on the retained proportion through central and tail-trimmed second moments, with finite-sample slack and complementary constructions establishing limiting tightness. We extend the results to sub-Gaussian vectors under additional distributional assumptions and obtain lower bounds for geometrically strong-mixing Gaussian sequences. For -subspace clustering under a low-rank Gaussian mixture model, we show that the clustering error of every global minimizer decreases polynomially with the signal-to-noise ratio under explicit sample-size conditions. For Gaussian robust regression, our analysis strengthens the uniformity of covariance and gradient-moment guarantees while improving the sufficient sample bound from to at constant confidence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.