Beyond Nominal Width: Comparison-Noise Geometry of Model Validation under Cross-Sectional Dependence
Abstract
Automated model discovery makes candidate generation cheap, but comparing candidates requires more than a large validation set. Counting evaluation units ignores both their dependence and how shared fluctuations affect differences between candidates. We introduce comparison effective width to characterize this comparison-noise geometry. The measure pairs candidate-centered score geometry with dependence among evaluation units. A projection-sufficiency result establishes that comparison rules invariant to common additive shifts depend only on centered candidate evaluation scores. Exact identities link comparison effective width to a comparison design effect and a second-order noise descriptor, all computable before reading evaluation outcomes. In label-free measurements across 2,258 matched cells from CSI300, a large-cap Chinese equity universe, as nominal width grows from 30 to 120, comparison effective width relative to nominal width falls from 0.726 to 0.389. The cell-level per-added-asset slope declines in 2,258/2,258 matched cells. CSI500, a disjoint universe in the same market, shows the same descriptive pattern with weaker attenuation. A same-candidate decomposition locates the CSI500–CSI300 gap in log comparison effective width mainly in return-side dimensionality and score–return alignment. Controlled synthetic experiments further show that the eigenvalue spectrum of the full covariance matrix alone does not determine split-half ranking agreement. An empirical relation combining cross-half candidate-IC covariance with zero-signal split-half IC-difference dispersion predicts this agreement on held-out synthetic worlds with cross-fitted R² ≈ 0.88 and systematic calibration deviations. Together, these results show why nominal evaluation counts should be complemented by a measure that accounts for evaluation-unit dependence and candidate-score geometry.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.