CoCurve: Cross-Module Co-Pruning Curvature for Structured LLM Pruning
Abstract
Resource-constrained deployment requires sustaining large language model (LLM) capabilities as models scale under fixed memory and computation budgets. Structured pruning advances this deployment frontier with smaller dense checkpoints, yet deciding what to prune remains bottlenecked by interactions among jointly removed units. We introduce Cross-Module Co-Pruning Curvature (CoCurve), which formulates structured pruning as set-dependent predictive risk over a unified inventory of attention heads and feed-forward groups. A co-pruning graph built from single-unit forward ablations assigns individual risk to nodes and reinforcement or cancellation to interaction edges, conditioning each decision on the units already removed. We evaluate 6 LLMs (3B–70B) and 3 vision–language models (VLMs) across 3 perplexity corpora, 12 language tasks, and 7 multimodal benchmarks. Across the five-model 20–40% grid and the 70B 10–50% sweep, CoCurve ranks first in 53/60 corpus comparisons (15/15 at 70B) and permits 2.2–6.6 more pruning points at matched quality in 9/10 interpolated comparisons, while leading all 6 VLM Avg7 blocks. After lightweight recovery, its retained structures remain strongest through 50% pruning, where the 8B checkpoint regains 10.8 Avg12 points; physical slicing delivers 1.58× dense prefill throughput with 41% lower peak memory. Mechanism analysis across 10 LLMs and 7 VLMs further finds organized within- and cross-module edge structure; matched removals of low-saliency, high-coupling units degrade 19/20 capability groups by up to 23.7 points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.