acceptodds
Under review as a conference paper at ICLR 2027

HarnessSHAP: Explaining LLM Agent Trajectories with Task-Adaptive Shapley Attribution

Abstract

When a large language model (LLM) agent succeeds or fails on a multi-step task, developers need to identify which parts of the trajectory contributed to the outcome. Existing attribution methods often rely on fixed judgment prompts, which may miss task-specific success criteria, and naive suppression can leave downstream actions dependent on removed steps. We introduce HarnessSHAP, a trajectory attribution method that learns a task-adaptive harness for a frozen LLM judge. The harness contains role-analysis and counterfactual-judgment prompts that estimate outcomes after dependency-consistent suppression while allowing downstream replanning. HarnessSHAP then applies Shapley attribution over dependency-valid orderings to assign contributions to evidence, reasoning steps, tool calls, and actions. The harness is optimized by XAI-Grad, a textual prompt optimizer that uses replay-derived endpoint and marginal errors to update the prompt pair without changing model parameters. We evaluate HarnessSHAP against a Fixed harness and Generic TextGrad on API-Bank, HotpotQA, ToolSandbox, and Who&When. XAI-Grad improves counterfactual outcome agreement and attribution quality across evaluated settings. On API-Bank, HotpotQA, and ToolSandbox, XAI-Grad achieves Spearman correlations of 0.74, 0.72, and 0.79, with endpoint MAE values of 0.103, 0.111, and 0.096, respectively. The ablation study shows that combining endpoint and marginal feedback reduces attribution error to 0.124. On Who&When, HarnessSHAP evaluates static failure localization separately from replay-based attribution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.