Causal Skill Attribution and Minimal Repair in Stochastic LLM Agents
Abstract
Auditing large language model agents requires identifying where decision disparities arise in execution traces, which executable components causally produce them, and how to repair those components while preserving utility. Stochastic trajectories, correlated activations, distributed interactions, and fine-grained carriers such as rule clauses, demonstrations, retrieval triggers, and routing edges can cause observationally suspicious components to be noncausal. We introduce Paired Factorial Skill Intervention and Repair (PFSIR), a framework for causal attribution and minimal repair over version-frozen, instrumented agent graphs. PFSIR applies type-preserving neutral replacements through balanced multi-component interventions on matched task pairs. It uses common random numbers and cached tool responses to reduce unrelated execution variation. A hierarchical sparse effect model estimates individual component effects and graph-constrained second-order interactions, and cross-fitting separates candidate selection from out-of-fold assessment. Separate models estimate disparity reduction and utility loss, enabling confidence-constrained optimization of the smallest patch that satisfies prespecified requirements for fairness, utility, regression risk, token changes, and execution cost. Fresh executions on isolated repair and regression sets validate candidate repairs, and the framework abstains when it identifies no acceptable patch. We evaluate PFSIR in three controlled agent environments covering resource matching, candidate ranking, and service routing. These environments contain known causal carriers, interaction and redundancy structures, correlated noncausal bystanders, and unbiased controls. We assess node- and edge-level recovery, disparity reduction, utility preservation, regression behavior, and structural generalization, yielding a closed-loop evaluation of intervention-defined attribution and executable repair under explicit replacement boundaries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.