When is predicting training loss worth the cost of initialization selection?
Abstract
Selecting among random initializations consumes computation that could instead train the selected network. We study this tradeoff under a fixed total compute budget using two predictors of future training loss: a separate frozen empirical neural tangent kernel (NTK) for each candidate, and a shared infinite-width limiting NTK used with each candidate’s own initial residual. For two-hidden-layer ReLU networks trained by full-batch gradient descent, the expected squared discrepancy between frozen-NTK predictions and actual outputs vanishes with width, uniformly over a fixed number of candidates and updates. We derive the limiting expected training loss after selection. Assume there are at least two candidates and the limiting initial-output covariance is positive definite. We compare selection by predicted future loss with selection by initial loss at the same training horizon. Future-loss selection yields a strictly lower limiting expected training loss if and only if the limiting dynamics scale residual directions by different absolute factors. Under a compute budget, this gain must also offset the updates forgone in scoring; we give an explicit sufficient condition in a realizable class with two error-decay rates. Extensive simulations illustrate the theoretical results and assess the tradeoff between prediction accuracy and computational cost..
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.