acceptodds
Under review as a conference paper at ICLR 2027

ASEMbed: Training-Free Embedding Extraction for Order-Robust Retrieval via Internal LLM Signals

Abstract

Training-free embedding methods reuse frozen decoder-only language models for retrieval without parameter updates. However, query content reordering can alter retrieval rankings and degrade retrieval performance. We observe that local attention concentrations remain visible under reordering, while their strength varies across layers. Based on these observations and prior findings linking attention-sink formation to representational compression, we propose ASEmbed, a training-free framework that uses Attention Sink and Projected Value signals to mitigate retrieval degradation under query content reordering. The framework comprises Boundary and Layer. Guided by these signals, Boundary seeks to uncover intrinsic interaction patterns that persist under reordering and uses them to partition each input into contiguous blocks. AS-Attention enables bidirectional attention across these blocks while retaining direct causal attention within each block. Layer traces the evolution of the same signals across layers to determine when to enable bidirectional cross-block interaction, preserving contextual information in earlier layers. It then selects readout layers from the representations computed with AS-Attention. Extensive experiments on three retrieval benchmarks and two reordering benchmarks demonstrate that ASEmbed achieves average nDCG@10 improvements of 1.96 and 4.28 points over the best-performing baseline on retrieval and reordering retrieval tasks, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.