BMIPrune: Class- and Layer-aware One-shot Average Activation-based Pruning for Convolutional Neural Networks via Mutual Information
Abstract
Activation-based pruning of convolutional networks has been pursued from two directions: class-agnostic criteria rank filters by the magnitude of their responses, while class-aware criteria rank filters by how strongly a response discriminates between classes. Both are applied across the network without regard to the depth of a filter’s layer, even though a filter’s role changes with depth. Early convolutions respond to generic low-level structure that is present regardless of the class, while later convolutions compose that structure into class-specific parts. A criterion that asks only what a filter reveals about the label therefore has little to mea- sure where filters are not class-selective, and a purely magnitude-based criterion may discard class information carried by a less active filter. This work proposes BMIPrune, a hybrid, one-shot criterion in which each layer’s own profile sets the balance between the two signals. Spatially averaged filter responses are modeled as class-conditional Gaussians, from which a Bayes mutual information (BMI) is computed for each filter. Weighting the posterior entropy by activation mag- nitude decomposes the resulting score exactly into a magnitude component and a class-information component, and their ratio defines each layer’s profile. On VGG16, the magnitude component accounts for about 98% of the profile in the first convolutions and as little as 68% in the last block, so the criterion falls back to magnitude where discriminative signal is absent and becomes class-aware where it is present. Before pruning, each layer is narrowed in isolation and its error curve recorded, giving a threshold-free redundancy score that is combined with the parameter cost of removing one filter from that layer. This sets each layer’s budget without reordering filters within the layer, so criterion and budget compose rather than interfere. On VGG19 with CIFAR-10, BMIPrune loses only 1.2 percentage points of accuracy at 60% parameter removal without fine-tuning, in line with other activation-based methods and more than 24 points ahead of uniform baselines. At 80% removal, where the baselines collapse, the loss is 7.7 points, 13 points less than the best competing activation-based method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.