acceptodds
Under review as a conference paper at ICLR 2027

Causal Skill Attribution and Minimal Repair in Stochastic LLM Agents

Abstract

Auditing large language model agents requires identifying where decision disparities arise in execution traces, which executable components causally produce them, and how to repair those components while preserving utility. Stochastic trajectories, correlated activations, distributed interactions, and fine-grained carriers such as rule clauses, demonstrations, retrieval triggers, and routing edges can cause observationally suspicious components to be noncausal. We introduce Paired Factorial Skill Intervention and Repair (PFSIR), a framework for causal attribution and minimal repair over version-frozen, instrumented agent graphs. PFSIR applies type-preserving neutral replacements through balanced multi-component interventions on matched task pairs. It uses common random numbers and cached tool responses to reduce unrelated execution variation. A hierarchical sparse effect model estimates individual component effects and graph-constrained second-order interactions, and cross-fitting separates candidate selection from out-of-fold assessment. Separate models estimate disparity reduction and utility loss, enabling confidence-constrained optimization of the smallest patch that satisfies prespecified requirements for fairness, utility, regression risk, token changes, and execution cost. Fresh executions on isolated repair and regression sets validate candidate repairs, and the framework abstains when it identifies no acceptable patch. We evaluate PFSIR in three controlled agent environments covering resource matching, candidate ranking, and service routing. These environments contain known causal carriers, interaction and redundancy structures, correlated noncausal bystanders, and unbiased controls. We assess node- and edge-level recovery, disparity reduction, utility preservation, regression behavior, and structural generalization, yielding a closed-loop evaluation of intervention-defined attribution and executable repair under explicit replacement boundaries.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.