Adversarial Bottleneck Explanations for Single-Pair CLIP-like Vision-Language Models
Abstract
Explaining a frozen vision–language model at the level of one image–text pair requires attributing a matching score without retraining or auxiliary examples. We study a bottleneck approach that separates feature selection from attribution readout. The Adversarial Bottleneck Method (ABM) searches over bounded intermediate-feature gates using the paired cosine score and assigns importance from a final-state Gaussian KL score. This removes the explicit compression–relevance coefficient used by M2IB while retaining a configurable perturbation path. We derive a finite-step bound for the projected gate formulation that accounts for boundary clipping and provides an explicit sufficient condition for one-step surrogate ascent. The bound does not imply optimal compression or attribution faithfulness. Reported experiments on Conceptual Captions, ImageNet, and Flickr8k show favorable image-side Confidence Drop point estimates under the inherited evaluation protocol, with complementary ROAD, component, and AltCLIP results. The submitted text wrapper has unresolved validity defects, and the historical measurements lack an exact run-to-implementation binding; they do not yet validate the reference formulation. ABM is a method for inspecting a model's matching score, with theoretical support for its update geometry rather than for causal or human-semantic explanations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.