acceptodds
Under review as a conference paper at ICLR 2027

QK-Wanda: Coupling Queries and Keys for Unstructured Pruning

Abstract

Wanda (Sun et al., 2024) prunes large language models by scoring each weight by its magnitude times its input activation, treating every linear projection independently. Self-attention violates this assumption: queries and keys enter the model only through their product, so the importance of a query weight depends on the keys, and vice versa. We introduce QK-Wanda, an unstructured pruning method that scores query and key weights by how much their removal perturbs the joint query-key product rather than either projection alone. Starting from a reconstruction loss on the causally masked query-key product, we derive a closed-form importance score for deleting a single weight; the result is Wanda's magnitude-activation criterion augmented with a term from the opposite projection. Scores are computed from a small calibration set without gradients or retraining, adding only 31-37% to Wanda's calibration and scoring time. On the Llama 3 family, QK-Wanda consistently reduces query-key reconstruction error and lowers perplexity relative to Wanda in most model-sparsity settings, with the largest gains at high sparsity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.