Beyond Best-Prompt AUROC: Textual Reachability of Distribution Shifts in VLM
Abstract
Models that connect images and text are increasingly used to describe visual shifts between image collections. A common approach is based on testing many textual descriptions and selecting the one that most strongly separates the two collections, as measured by Prompt AUROC. Yet, strong separation does not necessarily imply a meaningful explanation of the shift. In our experiments, the prompt "necessity", unrelated to blur, almost perfectly separates original images from blurred versions of them (Prompt AUROC 0.999). Our analysis explains how a large visual shift and a search over many descriptions can produce this result. Under a Gaussian model, we derive that Prompt AUROC depends on both the magnitude of the visual shift and how well text aligns with it. We show that we can largely predict the highest Prompt AUROC from the magnitude of the visual shift and the directional coverage of the candidate text embeddings, independently of the prompts' meanings. We thus introduce textual reachability, a label-free measure of the separation achievable through text relative to the best linear separation available from image embeddings. Across 113 diverse shifts, it distinguishes those captured by text from those where arbitrary words achieve similarly high AUROC. It also predicts 1) whether discovered words keep their alignment for other VLM backbones and 2) the performance gains from using text to adapt models to new image collections. Best-prompt AUROC correlates negatively with both.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.