Not Just What It Is, But What It Isn't: Comparative Meta In-Context Learning for Fine-Grained Categorization
Abstract
Visual in-context learning enables pretrained multimodal large language models to perform downstream categorization from labeled demonstrations without further parameter updates. Fine-grained categorization requires recognizing not only what a query is, but also what it is not among visually similar alternatives. While substantial effort has focused on demonstration retrieval and context construction, strengthening this comparative use of context through limited-data meta-training remains less explored. Standard cross-entropy (CE) supervises the correct label but does not explicitly penalize incorrect continuations once generation departs from the target sequence. Our diagnostic analysis shows that CE tuning raises both target and distractor scores, partially offsetting gains in the discriminative margin. We propose CLOUD (Contrastive Likelihood Only Upon Divergence), which complements target-label learning with explicit negative supervision on in-context distractor labels. Across three backbone configurations, CLOUD improves accuracy over CE-based Meta-ICL by 1.74–2.21% on six held-out natural fine-grained benchmarks and 2.14–6.11% on three specialized benchmarks. With CLOUD, Qwen3-VL-2B further outperforms Qwen3-VL-32B with vanilla ICL by 4.63% on average across all nine benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.