SparseForge: Efficient Semi-Structured LLM Sparsification via Annealing of Hessian-Guided Soft-Mask
Abstract
Semi-structured sparsity provides a practical path to accelerate large language models (LLMs) with native hardware support, but enforcing grouped sparsity constraints can substantially degrade model quality. Recovering this loss through extended sparse retraining incurs significant computational cost, motivating methods that achieve better recovery with fewer retraining tokens. We propose SparseForge, a post-training framework that jointly optimizes model weights and sparsity masks rather than relying solely on longer retraining. SparseForge combines Hessian-aware groupwise importance estimation with progressive soft-mask annealing, allowing the sparse support to adapt during recovery before final projection into a binary 2:4 pattern. On LLaMA-2-7B, SparseForge achieves 57.27% average zero-shot accuracy on the seven-task CAST-7 benchmark using 5B retraining tokens. Compared with prior SOTA, CAST, which achieves 54.37% with 7.86B tokens, this represents a 2.90% accuracy improvement with approximately 36.4% fewer retraining tokens. Evaluations across multiple model families further examine the applicability of the proposed recovery framework. These results highlight curvature-guided mask optimization and progressive hardening as a promising route to improving the accuracy–retraining-token trade-off. Our code is available at https://anonymous.4open.science/r/SparseForge-Anonymous.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.