Don’t Just Prune-and-Fine-Tune: Optimal Information Migration
Abstract
Structured pruning typically follows a simple recipe: rank, remove, and fine-tune. We argue that this conflates structural removal with functional discard. A structural unit may no longer warrant an independent place in the compressed network, while its learned contribution need not be discarded with it. Two compressed models can therefore have the same final architecture yet retain very different amounts of pretrained functionality, depending on how these contributions are handled before deletion. We reformulate structured pruning as a migration-before-deletion process and propose Optimal Information Migration (OIM). Information Preservation Priority (IPP) selects retained targets, while OIM uses optimal transport to migrate pretrained contributions onto those targets before deletion. We further introduce DeepInversion-based BatchNorm recalibration (DI-BN) to correct compression-induced statistical mismatch without real training data or updates to learnable parameters. Across CNNs and a vision Transformer on CIFAR-10/100 and ImageNet-1K, OIM consistently preserves substantially more pre-fine-tuning accuracy than direct pruning at the same compression budget and final architecture. On ResNet-50/ImageNet-1K at 30.26% FLOPs reduction, Top-1 accuracy improves from 15.79% after direct pruning to 42.15% with OIM and further to 52.55% with DI-BN.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.