acceptodds
Under review as a conference paper at ICLR 2027

SPELUNKA: A Framework for Finding Unanticipated Associations in Vision-Language Models

Abstract

Vision-Language Models (VLMs) are the backbone of many state-of-the-art multimodal systems, but often encode spurious and harmful associations from their training data. While prior work focuses on removing known biases from VLMs, methods for discovering unanticipated associations for retrieval tasks remain underexplored. We introduce \methodname, an automated, inference-time framework for finding unanticipated distributional biases for retrieval tasks in VLMs without pre-specified labeled subgroups, inference-time human supervision, or model fine-tuning. Given a concept of interest, \methodname generates captions for related images, extracts relevant entities from captions, and selects semantically meaningful associations. A frequency-based detection rule then identifies spurious associations by measuring over- or under-representation of these attributes in a VLM's top-ranked outputs relative to their full dataset distribution. We evaluate \methodname across four datasets, and find that it 1) successfully recovers both synthetic and naturally occurring spurious correlations, 2) achieves an average AUROC of 0.93 across four datasets in identifying images strongly associated with a query, and 3) outperforms baselines in finding features that are strongly associated with a given query.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.