acceptodds
Under review as a conference paper at ICLR 2027

Depth-Dependent Selection Geometry for Certified Structured Sparsification

Abstract

Structured pruning deletes whole channels, attention heads, or neurons, and a line of work makes such deletions happen during ordinary training by writing each weight as a product of D factors shared across a channel or head, so that weight decay drives whole groups to zero (deep gating). None of these methods says whether a particular deletion was justified: at a group whose factors are all zero, the training objective is stationary whatever the data say. Its curvature, however, depends sharply on depth. With two factors, the curvature at a deleted group, in that group's own coordinates, still sees the data in exactly the shape of the group lasso's optimality condition, so for almost every initialization a deletion that survives to a convergent gradient-descent limit carries a certificate that costs one gradient to check. With three or more factors, the data drop out of that curvature and a deleted group is locally stable in its own coordinates, justified or not. Because a network trained at one depth can be rewritten exactly at another, Certified Progressive Gating decides at depth two, records which decisions the certificate supports, and compresses deeper. On CIFAR-10 and CIFAR-100, it has the highest mean accuracy of eleven evaluated methods in five of six settings; on CIFAR-10, it matches a tuned dependency-aware pruner on two thirds of that pruner's training budget; on BERT-base, it has the highest mean accuracy of six evaluated structured-pruning methods, within noise. To our knowledge, it is the only method that records which of its deletions satisfy their optimality condition.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.