Decision Equivalence Does Not Identify the Explanation: Surrogate Multiplicity in Cluster Explanations
Abstract
A common way to explain a clustering is to train a classifier on its labels and explain that surrogate with SHAP. How much does the resulting explanation depend on the classifier? Across six datasets and six model families, estimated attributions were less similar across families than across bootstrap refits of the same family. Under a shared probability-space KernelSHAP protocol, the median paired similarity gap was ; its sign persisted across tested estimator settings, while its magnitude changed. We show that this ambiguity can remain even when decisions agree at every input. An exact construction rescales all class scores by the same positive, input-dependent factor, preserving every decision while making a decision-irrelevant feature uniquely most important. The result depends on what is explained: centering and normalizing scores gives exact attribution invariance within the common positive affine family. Decision fidelity alone therefore does not identify a score-based explanation. We recommend reporting the surrogate set and fidelity criterion, explained target, value function and background, and explanation variation across the selected models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.