acceptodds
Under review as a conference paper at ICLR 2027

LILA: Data-Free Structured Pruning of LLMs via NMF Spectral Geometry

Abstract

Structured pruning of large language models (LLMs) has universally required calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (Latent-Informed Layer Analysis) demonstrates that this dependency is avoidable: each FFN weight matrix is factorized with Non-negative Matrix Factorization (NMF) to obtain a parts-based spectral representation, and neuron importance is scored via the Kolmogorov Smirnov (KS) distance between singular value distributions of the full and neuron-ablated NMF matrix, a closed-form criterion requiring no calibration corpus, forward pass, or gradient computation. Two data-free matrix variants (Abs, Split) generate pruning masks from weight spectral geometry alone; a calibration-guided variant (Act) is provided as an internal ablation. Evaluation spans 7 model families (up to 70B). Without fine-tuning, data-free LILA surpasses PruneNet (45M-parameter RL policy) by up to 2.12 pp in zero-shot accuracy on LLaMA-2-7B and outperforms WikiText-2-calibrated SliceGPT by up to 6.1 pp across all sparsity levels, confirming data-free performance is not an artifact of omitting activation statistics. After one LoRA RFT epoch (matched WikiText-2 protocol), LILA Act surpasses SliceGPT+RFT by 6.39 pp on average using 98.4% fewer calibration samples (16 vs 1,024) across LLaMA-2-7B and Phi-2. NTK analysis confirms a reduction in functional distortion versus random pruning, providing theoretical grounding for the spectral criterion.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.