acceptodds
Under review as a conference paper at ICLR 2027

The Feature-Space Alignment Hypothesis for Neural Network Sparsity

Abstract

Why does something as simple as magnitude pruning retain accuracy even when most weights are removed, and why does performance eventually collapse at higher sparsities? This paper shows that the answer may lie in the ability to keep intact the representation of the original dense weight matrix when viewed in feature space, despite losing many weights to pruning. We study this question in a controlled setting where a linear embedding (representing prior layers) mixes input features across coordinates before classification, enabling direct analysis of the network's effective feature-space weights. We find that accuracy collapse under standard pruning coincides with divergence of the feature-space representation from its pre-pruning values. We then construct an optimization oracle that, given access to the embedding, selects a new sparse weight matrix, independent of the original, by directly preserving the effective feature-space matrix induced by the dense model. We show that under the same retraining budget, the oracle recovers performance at sparsity levels where standard pruning techniques degrade sharply, with consistent behavior across synthetic logic and more realistic tasks. Furthermore, we show that one can achieve oracle-level pruning performance even when the true embedding is not known by using a sparse autoencoder to recover its pruning-relevant structure, providing a path toward more general feature-aware sparsification. These results suggest a mechanistic explanation of sparsity and a guiding framework for feature-aware improvements in pruning methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.