A Principled Decomposition of Explanation Failure in Routed Models
Abstract
Feature scores in routed models can disagree with the model's response even when every prediction term is exposed. We decompose this failure for frozen-route gradient-times-input scores under a declared replacement into reference-origin mismatch, finite-step expert approximation, and omitted routing response. The exact decomposition connects each term to a correction or control. Centering removes origin mismatch and affine experts remove expert approximation. Routing can still reverse the response without a support switch. Across ten tabular datasets, five seeds, and four routed predictors, centering increases joint donor-replacement probability drops in 199/200 runs. Complete differentiation improves numerical-field sign agreement in 377/400 checkpoint–reference pairs, but finite-step errors remain. This motivates routing-derivative priorities for budgeted endpoint checks. An original-author SENN reproduction on three datasets extends the coefficient–response distinction to native explanations. Complete derivatives improve response agreement, while a logit-aligned training control incurs a predictive cost. The framework identifies which failures each correction addresses and why local agreement does not guarantee finite-response fidelity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.