acceptodds
Under review as a conference paper at ICLR 2027

SPD: Single Pass Decoding for Generative Reranking

Abstract

Large language models (LLMs) achieve state-of-the-art generative ranking quality, but the ranking they produce must be decoded, and autoregressive decoding spends one sequential forward pass per emitted token. We observe that the only tokens a ranker must emit are the N ordinal values, corresponding to the number of positions available to rank, naming the items in ranked order, and that this narrow, permutation-structured output format admits decoding strategies which are much more efficient than left-to-right generation. We introduce SPD (Single Pass Decoding), a format-specialized decoding strategy that decodes all N ordinals in O(1) forward passes. SPD reads an N × K item-position score matrix off the LLM’s prefill hidden states with a lightweight self-attention head, then decodes the ordinals as the optimal bipartite assignment of that matrix via the Hungarian algorithm, yielding a valid permutation by construction rather than by repair, without enumerating candidate prefixes or re-querying the LLM. Through a systematic study of training signals and backbone adaptation, we show that LoRA-based finetuning combined with autoregressive LLM ranking distillation reaches 28 ms endto-end inference, a speed-up of 64× end-to-end over the reasoning teacher while maintaining ranking quality on par. We provide a complete ablation decomposing the contributions of architecture, training signal, and backbone adaptation. Our framework connects generative ranking to combinatorial optimization, opening a path toward other O(1)-decode mechanisms for real-time ranking.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.