acceptodds
Under review as a conference paper at ICLR 2027

CiteProbe: From LLM Attention to Citations

Abstract

Sentence-level citations make Retrieval-Augmented Generation (RAG) outputs easy to verify by identifying the source sentences that support a generated answer. Yet existing methods face a trade-off between effectiveness and efficiency. Perturbation-based approaches require repeated LLM inference, while attention-based approaches typically rely on fixed aggregation rules that obscure informative variations across layers and heads. Can accurate citations instead be directly decoded from the attention weights without any additional pass? We show that they can, and that signals are concentrated in a few heads. To leverage this, we introduce CiteProbe, a lightweight approach that formulates sentence-level citation as a readout problem over the attention weights of a frozen LLM. For each candidate sentence, CiteProbe constructs an attention profile capturing the attention it receives from response tokens across individual layers and heads, thereby preserving the full layer–head structure. A lightweight, model-specific linear probe then learns to map these attention profiles to citation scores. This formulation exploits citation-relevant signals already present in the attention weights and scores all candidate sentences using a single generation pass, avoiding the repeated LLM inference required by perturbation-based methods. Experiments across six LLMs and four datasets show that CiteProbe consistently outperforms existing attribution methods while substantially reducing inference cost. These results demonstrate that sentence-level citation signals are directly decodable from structured LLM attention, offering an effective and efficient approach to citation attribution. We also show that CiteProbe can be applied to black-box LLMs under surrogacy mode. We release CiteProbe as open source at https://www.github.com/citeprobe/citeprobe.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.