Auditing Higher-Order Feature Interactions via Interventional Decomposition
Abstract
Signed pairwise interaction scores are widely used in explainable AI to classify feature cooperation as synergy (if positive) or redundancy (if negative). However, when interactions involve more than two features, pairwise scores fail: they either vanish entirely or flip sign across inputs, mislabeling synergy as redundancy. A single scalar summary is fundamentally insufficient to disentangle uniqueness (U), redundancy (R), and synergy (S). To overcome this limitation, we introduce InDe-X: Interventional Decomposition for eXplainability. While classical predictability decompositions study theoretical data capacity by retraining models on feature subsets, our method audits the internal causal pathways of a fixed, deployed network via post-hoc interventional masked inference, decomposing marginal predictive gains into per-feature U/R/S profiles over the Boolean lattice. We provide finite-sample concentration bounds, strict variance reduction via coupled diamond sampling, and uniform convergence guarantees over finite coalition vocabularies. InDe-X recovers ground-truth roles on tabular SCMs (overcoming up to 411x baseline conflation), surfaces synergistic and backup circuits in GPT-2 missed by activation patching, and identifies multi-region anatomical co-dependencies on ChestX-ray14 with superior deletion faithfulness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.