Decoupling Architecture Search Cost from Deployment Multiplicity
Abstract
Neural architecture search typically ranks candidates by the validation score of a single trained model, while some deployed systems use a multi-instantiation protocol in which an architecture is trained independently times and predictions are averaged. The risk of the deployed object is then governed jointly by single-model accuracy, predictive variance, and the correlation between replicas, and rebuilding the full -replica protocol for every proposal makes screening cost scale with . Under homogeneous regeneration and squared loss, we prove an exact reduction: deployment risk equals , so deployment-level ranking is determined by three architecture-level statistics that can be estimated from a fixed number of proxy instantiations, independently of the target . We then lift this population identity to finite validation data, deriving screening rules that control harmful acceptance with high probability when the stated validation-noise and proxy-stability bounds hold. We also delimit the result: under heterogeneous augmentation, whether a candidate helps depends on its error alignment with the deployed system and cannot be decided from candidate-only statistics. On Criteo-1M, Avazu-1M, and MovieLens-1M, the reduced objective attains Spearman – against measured deployment risk, versus – for single-run validation. Confidence-adjusted screening reduces the observed false acceptance rate from to on Criteo-1M and admits no harmful updates in the reported end-to-end runs, while requiring a fixed number of proxy trainings per candidate rather than full -replica reconstruction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.