acceptodds
Under review as a conference paper at ICLR 2027

Learning Verifiable Input Attribution in Language Models

Abstract

Language-model agents make decisions from context they did not write, such as retrieved passages, tool outputs and injected instructions. Input attribution estimates how much each input influenced a decision, which supports context management and helps audit actions steered by untrusted text. We define an input's influence counterfactually, as the change in the model's likelihood of its decision when that input is masked, that is, removed from the context. Leave-one-out (LOO) computes this exactly but needs one model call per input, while simply asking the model takes only a single call but is often unfaithful. We introduce the Decision Faithfulness Suite (DFS), which spans four decision tasks and labels every input with its exact LOO influence. Across the Qwen models and scales we test, prompted models find the inputs that support a decision better than they rank their influence, a pattern we call the support-influence gap. We propose Counterfactual Influence Distillation (CID), which trains the model with verifiable rewards to rank its inputs by influence and then measures only the top few exactly by masking them. We find that (i) CID is more faithful than the popular ContextCite on every task for both Qwen3.5-9B and Gemma-4-12B, with at least ten times fewer calls, (ii) prompted models also surpass ContextCite with enough masks, but training helps most when masks are scarce, and (iii) on controlled reasoning, a small trained model ranks influence better than prompted models eight times its size.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.