acceptodds
Under review as a conference paper at ICLR 2027

Identifying Support Knowledge Representation for Large Language Models

Abstract

Attention head pruning provides structured compression for large language models, yet identifying which heads carry target-domain capabilities remains challenging. To initiate the exploration of target-domain knowledge, we define the Support Knowledge (SK), which is the learned patterns of information selection and integration that support target-domain predictions. Based on the definition of SK, we further propose the Activation-Calibrated Margin (ACM) to determine their retention priority. The ACM combines intra-layer rankings of the token competition margin, defined as the difference between the top two attention logits, and the head output magnitude. During scoring, both statistics are collected in a single forward pass over calibration data without training or gradient computation. We then probe expanding candidate sets per layer, allocating a fixed pruning budget across layers based on marginal increases in loss and prediction disagreement proportion for each additional candidate. On domain datasets, ACM better maintains language modeling quality and prediction agreement with the full model than existing baselines, especially under higher pruning ratios, while achieving the highest accuracy in domain question-answering experiments among compared pruning methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.