FetNAS: Joint Optimization of Feature Selection and Neural Architecture Search
Abstract
Deep learning is increasingly deployed in resource-constrained settings such as IoT and industrial sensing, where memory and compute are limited. In these settings, input features often correspond to physical sensors that are costly to install and read. Because feature relevance is rarely known a priori and datasets often contain irrelevant or redundant variables, an effective model must select as few input features as necessary while also fitting a constrained memory footprint. Existing methods address only one side of this problem: Neural Architecture Search (NAS) treats the input features as fixed, whereas embedded feature selection assumes a fixed architecture, leaving the joint search space largely unexplored. We address this gap with FetNAS, a weight-sharing NAS method that treats the number of input features as an elastic dimension alongside network depth and width. FetNAS trains a single over-parameterized supernet using progressive shrinking, where each subnetwork uses the top-ranked features according to a gradient-based importance ranking rather than a random ordering. Once the supernet is trained, subnetworks are evaluated without retraining and a multi-objective evolutionary search identifies the Pareto front over predictive performance, parameter count and input dimensionality. Across the TabArena benchmark, FetNAS achieves a better performance-complexity trade-off than the baselines for classification. Up to the third quartile of model size, it typically uses - fewer parameters and - fewer input features. For regression, gains are more modest and budget-dependent. FetNAS improves performance at moderate feature budgets but retains a larger share of the features than for classification. Beyond reducing model complexity, the explicit feature subset lowers data-acquisition cost and eases edge deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.