Conceptdiff: Contrastive Concept Discovery for Dataset Comparison
Abstract
Identifying how two datasets differ in their semantics is a basic primitive for model understanding, data analysis, and scientific discovery. It is also hard to find: differences can be subtle, their form unknown in advance, and there are usually several. Current methods require foundation models to inspect the datasets, and propose natural-language hypotheses regarding the difference. This makes them expensive, and it limits them to differences the model thinks to propose. In this paper, we argue that differences should instead be learned directly from data. We formalize this as contrastive concept discovery (CCD), the task of finding a small set of dataset difference concepts that are discriminative (each separates the two datasets), monosemantic (each encodes a single factor), and diverse (together they cover distinct differences). We then introduce ConceptDiff, a lightweight, domain-agnostic method based on training a regularized concept-based classifier designed to learn diverse concepts. Across text, image, and video use-cases, we find that ConceptDiff recovers ground truth dataset differences better than LLM-based, clustering, and sparse-autoencoder baselines, and consistently finds concepts that more discriminative, diverse and monosemantic.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.