acceptodds
Under review as a conference paper at ICLR 2027

Shared-Response Agreement Fabrication: A Model-Side Attack on Multi-Explainer Visual Explanations

Abstract

Post-hoc visual explanations are often checked with multiple explainers. When several explainers highlight the same object region, this agreement is usually treated as evidence that the model relies on object semantics. We show that this interpretation can fail. We first place seven common explainers on a shared spatial partition and prove that their block-level heatmaps are driven by the same first-order deletion response, up to explainer-specific display maps and residual terms. Thus, multi-explainer agreement can reflect a shared local response rather than independent evidence. Based on this structure, we propose Shared-response Agreement Fabrication (SAF), a model-side attack that fabricates object-centered multi-explainer agreement. SAF trains the classifier so that a patch encodes the class decision, while a flat-germ objective suppresses the shared response on the patch-affected region and an object-region objective keeps it high on the original object region. After release, the classifier predicts from the patch, but the considered explainers still highlight the original object region. We provide a sufficient condition for this fabricated agreement and show that it does not imply the largest finite deletion effect. Experiments on standard image classification benchmarks show that SAF achieves higher attack success across datasets, model architectures, and explanation methods than representative explanation attacks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.