acceptodds
Under review as a conference paper at ICLR 2027

-Support Pruning: Tight Convex Relaxation and Principled Optimization for N:M LLM Pruning

Abstract

Semi-structured N:M sparsity—keeping N of every M weights, the form that supported sparse kernels can accelerate—is the dominant hardware-realizable way to prune large language models. Choosing which N weights survive is combinatorial, and current methods approximate that choice rather than characterize it: one-shot pruners commit to a ranking fixed before optimization, while learnable-mask methods optimize a nonconvex surrogate at the cost of gradients through the model. We show that this choice has a convex geometry of its own. Per block, the convex hull of the N-sparse unit Euclidean ball is exactly the unit ball of the k-support norm with k=N: the tight per-block relaxation and, among the classical structured penalties we consider, the one that carries the budget N itself; for N greater than or equal to 2, group and exclusive lasso do not. The relaxation is constructive. Its exact closed-form proximal operator is governed by one water level per block, found in a single sorted pass: weights above it survive and are shrunk, while the rest become exactly zero, followed by a least-squares refit that restores the survivors. It turns the layerwise pruning objective into a convex program solved in the one-shot post-training setting and certifies per block that rounding the relaxed solution back to N:M incurs no additional reconstruction loss. Magnitude and activation-scaled saliency reappear as zero-iteration gates, and under diagonal curvature with a separation condition, hardening and refitting the solution returns the combinatorial optimum. Read across matrices rather than within a block, the same water level prices the sparsity budget, so one convex object decides both which weights to keep and where sparsity is spent. Across seven models from four architecture families, KSP achieves the best accuracy in six of seven columns under 2:4 sparsity and the best results in eight of ten cells for each metric at 75% sparsity. On a 70B model, allocating the budget improves over uniform N:M sparsity at matched measured latency, and pricing the allocation by latency provides further gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.