acceptodds
Under review as a conference paper at ICLR 2027

PullSHAP: Principled Attribution Pullback Through Learned Representations

Abstract

Learned representations create an attribution mismatch: predictions are made after transforming the original features, while explanations are often required in the original feature space. We study when Shapley attribution can be pulled back through a factorization . We first show that no generic local chain rule exists: two smooth representations can agree on neighborhoods of both the baseline and explicand, and hence have identical endpoint derivatives of every finite order, yet induce different composite input-feature Shapley values. Moreover, representation compression alone is insufficient; even a smooth scalar representation can realize any normalized coalition game on the relevant hybrid vertices. We then establish sufficient conditions for pullability. Mixed cross-feature curvature of the representation controls its deviation from an additive coalition surrogate, showing that nonlinearity alone need not cause representation-side coalition distortion. On this surrogate, a distinct obstruction is the discrepancy between Aumann–Shapley path allocation and Shapley–Shubik allocation. These results lead to \PullSHAP, which combines integrated representation contributions with integrated downstream sensitivity, together with an error decomposition separating representation distortion, downstream allocation mismatch, and numerical approximation. Controlled experiments confirm the predicted exact and failure regimes. Across 1,250 exact attribution instances nested within successful learned-representation runs, median relative- error is , with a non-negligible upper tail. Within dense MLP encoder tracks across three datasets, representation-side distortion is positively associated with exact attribution error. A separate exact-reference benchmark compares the differentiable approximation with established SHAP estimators while keeping black-box model-evaluation budgets distinct from gradient/Jacobian computation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.