acceptodds
Under review as a conference paper at ICLR 2027

No Man Is an Island: Overlapping Pruning for Near-Unstructured Quality at 2:4 Efficiency

Abstract

Semi-structured 2:4 sparsity is the dominant fine-grained pruning pattern on modern ML accelerators, yet for large language models it degrades quality far more than unstructured pruning at the same 50% sparsity. We trace this gap to mask rigidity: 2:4 partitions weights into disjoint four-element windows, so both selectors draw from the same four inputs, severely constraining the feasible mask space. We propose Overlapping Pruning (OP), an algorithm–hardware co-designed sparsity structure that overlaps neighboring selector windows, enlarging the mask space while preserving 2:4's selectors, 2-bit index, and datapath. At 4-, 8-, and 16-bit, OP-1D and OP-F2D cost at most 2:4's core area, - its delay, and - its power. OP-2D additionally yields transposable masks. Feasible masks are the independent sets of a transversal matroid, so a saliency-ordered greedy pass with a per-candidate feasibility check is exact for every OP pattern; our kernels generate OP-F2D masks within the 2:4 mask-generation cost. One-shot OP-F2D with Wanda++ increases perplexity by at most over unstructured 50%, versus up to for 2:4. OP-1D attains lower perplexity than 4:8 and OP-F2D stays within of 8:16; patterns cost - 2:4's core area and - its delay. OP-F2D matches or improves on reported results from learned-mask and factorized 2:4 methods at a fraction of their pruning cost, with no mask training or added parameters. By relaxing mask geometry alone, OP-F2D closes - of the unstructured-to-2:4 quality gap at 2:4's hardware cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.