acceptodds
Under review as a conference paper at ICLR 2027

Recovering Decision Provenance from Internal Activations in LLM Agents

Abstract

Tool-using language-model agents combine information from user instructions, policies, retrieved records, memories, and tool observations when producing structured actions. Although the resulting tool call records the selected tool and argument values, it does not reveal which parts of the context influenced each field. We define field-level decision provenance through source interventions: for a fixed proposed action, we neutralize one source at a time and measure how the model's score for each realized field value changes. This yields a signed effect matrix over sources and action fields that captures both the direction and strength of source influence, but computing it directly requires repeated model evaluations. We introduce the Interventional Effect Decoder (IED), a lightweight decoder trained on offline intervention targets. Given the original context and proposed action, IED reads internal activations from a frozen backbone and predicts the complete effect matrix from a single scoring pass, avoiding repeated interventions at deployment. Experiments across diverse tool-use settings show that IED recovers informative source-level provenance while substantially reducing online attribution cost. The predicted provenance also improves the repair of incorrect tool arguments by helping a reviewer locate the context relevant to each field. Finally, we extend this representation with support and authorization supervision in a separate decision model for prompt-injection defense, enabling pre-execution detection of actions influenced by sources that lack authority over the affected fields.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.