From Heatmaps to Reliance Fields: Diagnosing Visual Explanations
Abstract
Visual explanation methods are widely consumed under the assumption that spatial heatmaps directly expose model reliance. However, what an explanation highlights, what functionally supports a prediction under intervention, how evidence propagates across internal representations, and what is semantically meaningful are fundamentally non-interchangeable. We prove in Theorem 1 that finite-order local attribution signals do not uniquely determine finite-intervention reliance. This conflation obscures critical failure modes, including highlighted-but-unused evidence, hidden reliance, and potential shortcut cues. To overcome this limitation, we introduce Reliance Uncertainty Fields (RUF), which formulate latent functional reliance as a layer-wise posterior distribution over an Energy-Based Factor Graph. RUF couples local preference, finite-intervention influence, representation flow, stability, and optional semantic references into a unified probabilistic field. Building on RUF, we propose RIFT (Reliance Inference over Flow-based Traces), a diagnostic framework that computes posterior reliance means, marginal entropy (uncertainty), preference–function disagreement, flow attenuation, and shortcut candidates. RIFT reframes visual explanation from generating plausible heatmaps to inferring operational functional reliance under a specified intervention family and explicit uncertainty.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.