Condition on the Gallery, Not Just the Query: Query Equalization and Sparse Transport for Debiasing Vision-Language Retrieval
Abstract
Text-to-image retrieval with a vision-language model such as CLIP ranks the images of a gallery against a text query, and for a stereotyped query the retrieved set can have a demographic composition far from that of the gallery. A retrieval score has two arguments, the query and the image, yet many post-hoc corrections derive their edit from text prompts alone. We propose BALLAST (Barycentric Level-Allocated Latent Sparse Transport), a debiasing method for frozen CLIP models that conditions the correction on unlabeled images from the gallery's population. From prompts naming the protected attribute, BALLAST gives each reference image a soft membership over the attribute groups, and this one estimate drives two corrections: a repair of the gallery, which moves the groups' activations in the code of a frozen sparse autoencoder part of the way toward a shared barycenter on the coordinates where the groups differ, and a closed-form equalization of each query against the repaired gallery's group centroids. No per-image annotation is read and no gradient-based training takes place. On the SEM benchmark, with three retrieval datasets and four CLIP backbones, BALLAST lowers both KL divergence and MaxSkew below every SEM variant in all 24 comparisons on the two ViT backbones where it cuts KL on the stereotype queries by a factor of 2.3 to 5.0 relative to the best SEM variant per cell. The query equalization supplies most of the improvement, and its subspace has to come from the gallery: the same projection built from the attribute prompts alone raises KL above base CLIP on all 24 retrieval cells of the benchmark.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.