Gradient Difference Enables Efficient Second-Order Vision-Language Interaction Explanations
Abstract
Vision-language encoders such as CLIP excel at connecting visual and linguistic representations, yet interpreting their predictions remains challenging. Existing first-order explanation methods identify important image regions or text tokens individually, but fail to capture cross-modal feature interactions necessary for understanding dual-encoder predictions. Recent studies have therefore moved beyond first-order attribution toward second-order interaction explanations. However, the current frameworks face high computational or memory costs as text lengths and candidate interactions grow. To address these limitations, we propose a lightweight gradient-difference framework that derives second-order interaction explanations directly from existing first-order gradient-based attribution methods. By masking a target textual concept or visual region and measuring the resulting change in matching score gradients, our approach isolates phrase-conditioned visual explanations and region-conditioned textual explanations without introducing specialized interaction machinery or expensive sampling. Applied to standard methods (yielding Grad-CAM+, GAME+, and Grad-ECLIP+), our framework achieves performance comparable to state-of-the-art second-order approaches while significantly reducing computational overhead. The quantitative evaluations and human studies further demonstrate that our approach yields both faithful and user-informative vision-language correspondences.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.