acceptodds
Under review as a conference paper at ICLR 2027

Learning to Attribute by Training Internal Representations

Abstract

For any specified factual span in an answer generated by a large language model (LLM) with retrieval-augmented generation, fine-grained attribution aims to locate the fine-grained evidence in the context that best supports that span. This enables humans to efficiently audit LLM outputs and ensures the reliability of the results. Among approaches to this task, methods based on the model's internal representations offer fast inference, high accuracy, and pluggability. However, existing methods either apply a pretrained model directly for attribution or train only a small number of auxiliary parameters; no prior work has attempted to improve attribution by training the model's internal representations themselves. In this paper, we identify two obstacles during training for attribution: (1) the attribution-relevant training signal is sparse and correlated, while standard contrastive objectives can not focus on this sparse ranking signal, and (2) the value-vector cheating problem. To address these obstacles, we propose TIR (Training Internal Representations), which uses a margin ranking loss weighted by hard negatives and applies a stop-gradient to the value-norm component. Experiments show that TIR consistently improves attribution across multiple backbone models, multiple datasets, and multiple internal signals, and even exhibits a degree of cross-domain generalization, demonstrating the feasibility of obtaining more accurate attribution models through training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.