Classify by Descriptions: A Reality Check
Abstract
Classify by Descriptions (CbD) has been shown to both improve zero-shot image classification and provide natural language explanations of predictions. Instead of relying on costly per-task human expert training data to achieve this, CbD infers per-class descriptors from large language models (LLMs) and then uses a vision-language model (VLM) to recognize their appearance in a query image. In this work, we critically analyze CbD and characterize its behavior when recognition becomes more fine-grained and target categories might be underrepresented in LLM and VLM pre-training data. Existing evaluation setups are unable to support this analysis with either limited coverage of realistic problem domains or a single granularity of classification. To address this gap, we introduce the classification benchmark TaxoBench, a benchmark spanning 7 realistic classification domains and multiple levels of classification granularity. Using TaxoBench, we find that CbD performance drops substantially as categories become more fine-grained and less familiar to the underlying VLM. We further identify two key sources of failure: LLM-generated descriptions often fail to capture the visual distinctions required by the task, and VLMs struggle to reliably recognize the resulting open-vocabulary concepts. We evaluate existing and new mitigation strategies but find only modest improvements, emphasizing the need for further research in more robust pre-training of foundational models and CbD as a whole.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.