SAM-RISE: Signed Black-Box Saliency for Vision-Language Models via Semantic Region Ablation
Abstract
The most capable vision-language models (VLMs) are closed and reachable only through APIs, which rules out interpretability tools that need gradients, attention maps, or activations. Black-box occlusion saliency, as in RISE and D-RISE, needs only the model’s output, but it does not transfer to VLMs unchanged: a VLM reads a tight occlusion mask as an object in the scene (masking a cat’s eyes yields “the cat has black eyes”), and probability-weighted masks saturate on a confident model. We introduce SAM-RISE, a black-box saliency method built around both failures. It removes a few dozen semantic regions from Segment Anything, each as a grey bounding box, scores the change in the unsaturated logit margin, and keeps both signs: positive evidence, regions whose removal lowers confidence in the answer, and negative evidence, regions whose removal raises it. When an API returns no token probabilities, sampled answers take their place. Semantic regions pay off in three ways. At the small query budgets a paid API imposes, they localise objects as well as random masks with 2 to 4× fewer queries. Negative evidence is actionable: subtracting it corrects over-confident “yes” answers on NaturalBench for two model families. And because the same part can be matched across images, per-image maps aggregate into class-level maps that show whether one image’s explanation is typical of its class, and that separate classes the model recognises from pixels from those it answers from language priors. We evaluate on MS-COCO, ImageNet, and NaturalBench, compare against white-box Grad-CAM and attention rollout and against human judgments, and run SAM-RISE on the closed Gemini model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.