acceptodds
Under review as a conference paper at ICLR 2027

Importance Is Not Gauge-Invariant: Measuring, Not Estimating, Unit Redundancy in LLMs

Abstract

Structured pruning ranks attention groups and FFN channels by importance proxies such as weight magnitude or activation traffic. We show these proxies do not measure the network's function: residual transformers admit per-group *gauge freedoms* — function-preserving rescalings of , , and FFN pairs — under which magnitude-based rankings reorder almost arbitrarily (Spearman –) while no output bit changes (max logit difference ). Interventional ablation rankings are exactly invariant under the bit-exact gauge on Llama-3.1-8B (rank survival ); on QK-normalized architectures the gauge is only -approximate, and the observed rank change (Qwen3: ) tracks the induced function drift. Across six backbones from five institutions, estimation-based proxies correlate with interventional importance at ; *exactly* invariant proxies (Taylor, Fisher) do no better (; a Taylor top- keep-set yields the NLL of random, vs. ) — invariance is necessary, not sufficient. A pre-registered mechanism hypothesis (QK-norm predicts proxy sign) is falsified by our own test, and gauge canonicalization does not rescue the proxies. Measurement has pitfalls of its own: on Mistral-7B's flat redundancy landscape, extreme solo-score selection is catastrophic (NLL vs. random) until one mid-budget score refresh restores safety. We therefore develop MAPS (Measured Allocation of Pruning Structure), a measurement-based pruner with a single primary hyperparameter that selects heterogeneous units (KV groups, FFN groups, whole layers) by *measured* candidate comparison. Without recovery training, at a nominal keep ratio of it retains 99.5% of zero-shot MC ( vs. ) and 99.8% of GSM8K on OLMo-3-32B. Across four backbones and three budgets, measurement-based selection tops every comparison: MAPS takes 10 of 12 backbone–budget cells (up to points over the best structured baseline), a fully-measured depth search (SLEB) takes the deepest hub cell — a dose–response in the amount of measurement — and estimation-based criteria (magnitude, Wanda-style, Taylor, FLAP-style) trail throughout, with score families backbone-dependent under a fixed allocator. An allocation-only ablation (global per-parameter greedy over the same scores, even with one mid-budget re-measurement) still trails random: *where* you remove matters as much as *what* you score, and *how much* you measure decides who wins.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.