Detecting Fine-Grained Differences Between Image Datasets
Abstract
Detecting whether and in particular _how_ two data distributions differ is a fundamental statistical problem. For example, we might be interested in which concepts (expressed in natural language) differ between two classes, the outputs of different generative models, or the same model with different prompt specifications. When working with large and complex image datasets, we intuitively think that first selecting concepts from a broad list using fine-grained patch embeddings, and then optimizing the concepts to better fit the data is an effective workflow. However, prior work instead uses MLLMs to propose concepts and image-level embeddings to select them. Therefore, we systematically compare these techniques and find that selecting concepts from broad lists can replace MLLM-proposals at obviously lower cost. MLLM inference can then be spent on optimizing selected concepts instead to achieve substantial improvements. Additionally, by designing a new benchmark that focuses on low-prevalence and multi-concept distinctions, we find that selecting nonredundant concepts instead of only the most distinguishing concepts matters, and we introduce techniques for how to effectively do this. Finally, we show how our methods address important problems like detecting spurious correlations and biases. Overall, our work presents a fundamental perspective shift on how we analyze interpretable distributional differences between datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.