acceptodds
Under review as a conference paper at ICLR 2027

Subspace Node Pruning: Retraining-Free Neural Network Compression via Orthogonal Subspace Projections

Abstract

Node pruning is the art of removing computational units such as neurons, filters, attention heads, or even entire layers to significantly reduce inference time while retaining network performance. In this work, we show that a broad family of importance measures, from activation-based to Fisher-based, share a common quadratic form, and that through a lower-triangular change of basis we can reach a pruning subspace in which whole nodes can be pruned with a simultaneous recovery of the impact of lost units via linear least squares. Furthermore, the importance measures within our subspace are decoupled and additive through an orthogonalization of the layer activations, allowing multi-node pruning within layers without interference between importance scores. Our method, Subspace Node Pruning (SNP), outperforms the compared baselines on VGG-16 and ResNet-50 on ImageNet. On OPT-125M and LLaMA-2 7B, it matches the state-of-the-art LLM Surgeon method at a 40-shot schedule and outperforms it in every case where fewer shots are used. All of this is achieved without any retraining after pruning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.