Tapering Neural Networks with Temporally Aggregated First-Order Importance
Abstract
In gradual pruning, network architecture is compressed simultaneously while learning the network parameters. Importance scoring of the parameters to prune allows to compare them such as for their ranking to cull a desired parameter set at a given pruning instance. In this paper we propose Temporal Aggregation of Parameter Effects for Reduction (TAPER), which uses an effective first-order estimation of weight importance on lost network capacity, combined by aggregation of importance scores for statistical robustness of gradual pruning decisions to batch variations. TAPER is evaluated favorably against the state of the art in gradual pruning in standard datasets using established metrics. It performs well both for pruning simpler CNNs such as VGG as well as more complex architectures such as ConvNeXt. Since network accuracy metrics compound major effects from training choices with pruning criteria, we also introduce Pruning Impact as a differential metric for assessing pruning criteria better isolated from cumulative training effects.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.