A Good Initialization is All You Need for Faithful Visual Attribution
Abstract
Faithful visual attribution, which aims to identify which image regions genuinely support a model's prediction, is a central challenge in interpretability research. Search-based perturbation methods define the insertion–deletion faithfulness Pareto frontier by masking regions and measuring score changes, yet they must construct a complete ordering of all regions, incurring substantial cost when only a compact top-\(k\) evidence mask is needed. We study this mask-first problem from two angles. An analysis of the behavioral gap between greedy and phase-window based search yields CoPAIR, a coarse pairwise initializer that already improves search. A deeper rethinking of the problem leads to TRACE (Top-\(k\) Region Attribution via Cross-Entropy), which directly optimizes a binary evidence mask via cross-entropy at a cost comparable to gradient-based methods. Both can serve as compact attribution masks or warm-start search methods when a complete ranking is required. Across ImageNet classification with CLIP ViT-L/14, CLIP RN101, and ResNet-101, our initialized search methods establish a new state-of-the-art frontier for faithful full-ordering attribution under inclusive forward-call accounting. On POPE and RePOPE with Qwen2.5-VL-3B-Instruct and LLaVA-v1.5-7B, TRACE+Greedy gives the strongest search-based MLLM attribution results. Direct TRACE masks further achieve single-point RePOPE repair rates of \(94.44%\) and \(96.00%\), showing that compact evidence masks can be actionable attribution outputs, not merely prefixes of full rankings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.