Identifying Reliable and Unreliable Components of Shapley Values
Abstract
We propose a method to decompose Shapley values into reliable and unreliable components. First, we prove that, given a set of AND–OR interaction patterns modeled by the DNN, evenly allocating the utility of each interaction to the input units involved in the interaction yields the Shapley value. Moreover, interaction patterns can be categorized into reliable and unreliable interactions according to their transferability across different DNNs. In this way, we define the Shapley attributions allocated from reliable interaction patterns as the reliable Shapley components, which reflect the relatively browntrustworthy decision logic learned by the DNN, while the Shapley attributions allocated from unreliable interaction patterns are termed unreliable Shapley components. We prove that the decomposed reliable and unreliable Shapley components still satisfy the extended four Shapley axioms. brownExperiments demonstrate that the proposed reliable and unreliable Shapley components capture decision logic with markedly different representational qualities: (1) Reliable Shapley components are substantially more similar across LLMs and more robust to input perturbations than unreliable Shapley components. (2) Unreliable Shapley components tend to exhibit greater cancellation between positive and negative attributions, thereby contributing less to the LLM’s predictions. (3) Reliable Shapley components tend to concentrate more heavily on input units relevant to the prediction than unreliable Shapley components.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.