When Are Feature Crosses Discoverable?
Abstract
Why can explicit feature crosses help models that can already represent interactions? Expressive capacity does not ensure that finite data suffice to fit and certify a useful additive correction to a trained predictor. We formalize this gap as checkpoint-local discoverability, separating residual fitting from post-selection certification. A local risk expansion connects one-step utility to curvature-weighted residual prediction. In a fixed-design Gaussian model, we derive a sufficient exposure–cost frontier for linear smoothers that balances accessible signal, noise fitting, and search complexity. For nested Gaussian projection learners that fully expose the same signal, refinement strictly lowers the probability of positive gain by adding noise-only directions. We introduce CROSSAUDIT, a sample-split audit that selects the pre-validation update with the largest positive simultaneous lower confidence bound and otherwise returns a null update. For bounded logit updates under logistic loss, every positive certificate guarantees exact conditional population-risk improvement at the prescribed confidence level despite candidate selection. Controlled synthetic and semi-synthetic CTR experiments reproduce the refinement effect: atomic representations can lose utility while preserving the target function, whereas hierarchical sharing retains the coarse signal. Real-label audits further show that pair-indexed utility need not be interaction-only: main-effect repairs can be stronger, while the tested post-fit residualized pairs receive no positive certificates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.