HarnessSHAP: Explaining LLM Agent Trajectories with Task-Adaptive Shapley Attribution
Abstract
When a large language model (LLM) agent succeeds or fails on a multi-step task, developers need to identify which parts of the trajectory contributed to the outcome. Existing attribution methods often rely on fixed judgment prompts, which may miss task-specific success criteria, and naive suppression can leave downstream actions dependent on removed steps. We introduce HarnessSHAP, a trajectory attribution method that learns a task-adaptive harness for a frozen LLM judge. The harness contains role-analysis and counterfactual-judgment prompts that estimate outcomes after dependency-consistent suppression while allowing downstream replanning. HarnessSHAP then applies Shapley attribution over dependency-valid orderings to assign contributions to evidence, reasoning steps, tool calls, and actions. The harness is optimized by XAI-Grad, a textual prompt optimizer that uses replay-derived endpoint and marginal errors to update the prompt pair without changing model parameters. We evaluate HarnessSHAP against a Fixed harness and Generic TextGrad on API-Bank, HotpotQA, ToolSandbox, and Who&When. XAI-Grad improves counterfactual outcome agreement and attribution quality across evaluated settings. On API-Bank, HotpotQA, and ToolSandbox, XAI-Grad achieves Spearman correlations of 0.74, 0.72, and 0.79, with endpoint MAE values of 0.103, 0.111, and 0.096, respectively. The ablation study shows that combining endpoint and marginal feedback reduces attribution error to 0.124. On Who&When, HarnessSHAP evaluates static failure localization separately from replay-based attribution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.