Sublinear Estimation of Additive Interpretability under SHAP
Abstract
Additive feature attribution methods explain model predictions as sums of per-feature contributions, but their fidelity is fundamentally limited by how well the model behavior can be represented additively. We formalize this limitation through , defined as the best achievable coefficient of determination over the class of additive functions. We reduce the problem to explained-variance estimation and prove that a sublinear learnability estimator can estimate additive interpretability under the Shapley coalition distribution without recovering the optimal attribution vector. For any fixed error tolerance and success probability, the estimator requires model evaluations in the high-dimensional regime. Across synthetic and real-world experiments, it accurately estimates additive interpretability with fewer model evaluations than attribution-based baselines. These results enable practitioners to assess whether an additive explanation can faithfully represent a prediction before paying the cost of computing its feature attributions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.