acceptodds
Under review as a conference paper at ICLR 2027

Decision Equivalence Does Not Identify the Explanation: Surrogate Multiplicity in Cluster Explanations

Abstract

A common way to explain a clustering is to train a classifier on its labels and explain that surrogate with SHAP. How much does the resulting explanation depend on the classifier? Across six datasets and six model families, estimated attributions were less similar across families than across bootstrap refits of the same family. Under a shared probability-space KernelSHAP protocol, the median paired similarity gap was ; its sign persisted across tested estimator settings, while its magnitude changed. We show that this ambiguity can remain even when decisions agree at every input. An exact construction rescales all class scores by the same positive, input-dependent factor, preserving every decision while making a decision-irrelevant feature uniquely most important. The result depends on what is explained: centering and normalizing scores gives exact attribution invariance within the common positive affine family. Decision fidelity alone therefore does not identify a score-based explanation. We recommend reporting the surrogate set and fidelity criterion, explained target, value function and background, and explanation variation across the selected models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.