Shared-Response Agreement Fabrication: A Model-Side Attack on Multi-Explainer Visual Explanations
Abstract
Post-hoc visual explanations are often checked with multiple explainers. When several explainers highlight the same object region, this agreement is usually treated as evidence that the model relies on object semantics. We show that this interpretation can fail. We first place seven common explainers on a shared spatial partition and prove that their block-level heatmaps are driven by the same first-order deletion response, up to explainer-specific display maps and residual terms. Thus, multi-explainer agreement can reflect a shared local response rather than independent evidence. Based on this structure, we propose Shared-response Agreement Fabrication (SAF), a model-side attack that fabricates object-centered multi-explainer agreement. SAF trains the classifier so that a patch encodes the class decision, while a flat-germ objective suppresses the shared response on the patch-affected region and an object-region objective keeps it high on the original object region. After release, the classifier predicts from the patch, but the considered explainers still highlight the original object region. We provide a sufficient condition for this fabricated agreement and show that it does not imply the largest finite deletion effect. Experiments on standard image classification benchmarks show that SAF achieves higher attack success across datasets, model architectures, and explanation methods than representative explanation attacks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.