acceptodds
Under review as a conference paper at ICLR 2027

The Geometry of Sparsity: Fisher Information Across Neural Network Pruning Methods

Abstract

Modern neural networks contain substantial parameter redundancy, yet the geometric consequences of removing parameters remain poorly understood. We study pruning through the lens of Fisher geometry and investigate how restricting a network to a sparse parameter subspace affects statistical dependencies among its remaining parameters. We introduce a per-dimension log-determinant ratio that quantifies deviation of the Fisher Information Matrix (FIM) from diagonality, providing a measure of weight-space coupling. We show theoretically that, under a fixed-Jacobian damped Gram-matrix surrogate, the pruning operation monotonically reduces this coupling. Empirically, across architectures, datasets, and pruning methods, we observe a consistent reduction in Fisher coupling over broad sparsity ranges, beyond what can be explained by parameter count alone. As the FIM becomes increasingly diagonal, updates produced by diagonal adaptive optimizers become more closely aligned with natural-gradient updates, whereas the same behavior is not observed for SGD. Together, these results suggest that pruning does more than reduce model size: it systematically simplifies the local geometry of the surviving parameter space, producing representations that are more coordinate-separable and better approximated by diagonal second-order structure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.