Better Heads, Better Rankers: Selecting and Training Attention Heads for Listwise Reranking
Abstract
Attention-based ranking has become a key paradigm for listwise ranking in large language models (LLMs). It reads the attention weights between internal query tokens and document tokens as the basis for ranking, thereby dispensing with autoregressive decoding and offering markedly better efficiency and interpretability than generative rankers. However, its effectiveness depends on two interrelated design choices: which attention heads supply the ranking signal, and how to perform training augmentation on these heads. Previous studies have not yet established a unified paradigm for these designs. In this paper, we conduct a systematic study of head selection and training for attention-based ranking. On the selection side, we show that prevailing practice of averaging per-head scores over a development set is brittle and the resulting set lacks overall representativeness. We instead propose an frequency-based criterion that can select adaptive, representative head sets. On the training side, we introduce a unified two-stage recipe that optimizes the selected heads. During supervised fine-tuning, att-sft augments the training objective with loss terms computed from the head-aggregated attention scores. During reinforcement learning, we first identify the discrepancies between training and inference, and then propose att-dpo, which can directly align preferences on the attention distribution. Across two backbones and several ranking benchmarks, the recipe consistently surpasses untrained and training-related baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.