Recovering Decision Provenance from Internal Activations in LLM Agents
Abstract
Tool-using language-model agents combine information from user instructions, policies, retrieved records, memories, and tool observations when producing structured actions. Although the resulting tool call records the selected tool and argument values, it does not reveal which parts of the context influenced each field. We define field-level decision provenance through source interventions: for a fixed proposed action, we neutralize one source at a time and measure how the model's score for each realized field value changes. This yields a signed effect matrix over sources and action fields that captures both the direction and strength of source influence, but computing it directly requires repeated model evaluations. We introduce the Interventional Effect Decoder (IED), a lightweight decoder trained on offline intervention targets. Given the original context and proposed action, IED reads internal activations from a frozen backbone and predicts the complete effect matrix from a single scoring pass, avoiding repeated interventions at deployment. Experiments across diverse tool-use settings show that IED recovers informative source-level provenance while substantially reducing online attribution cost. The predicted provenance also improves the repair of incorrect tool arguments by helping a reviewer locate the context relevant to each field. Finally, we extend this representation with support and authorization supervision in a separate decision model for prompt-injection defense, enabling pre-execution detection of actions influenced by sources that lack authority over the affected fields.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.