When Examples Fail to Teach: The Demonstration Gap in Vision-Language Models
Abstract
Multimodal in-context learning (ICL) can help or harm vision-language model predictions, complicating demonstration selection. To understand this difficulty, we quantify the structure of demonstration utility across 13 models, three multiple-choice visual question answering datasets, and 27 conditions through exhaustive one-shot query–demonstration evaluations, measuring changes in correct-answer support and prediction correctness. Models that solve many of the same questions at zero-shot often benefit from different demonstrations on questions they both get wrong, despite partial utility transfer. On zero-shot errors, demonstration main effects are small, while query and pairing effects dominate utility variation. Selecting demonstrations by their average rescue rates rarely yields consistent gains across held-out query splits. Measured content similarities show weak linear alignment with pair-specific rescue patterns. In two end-to-end settings, the evaluated retrievers do not consistently improve accuracy. As a simple example of using query characteristics, we combine a simple trained gate for likely errors with voting over separate one-shot predictions using random demonstrations. This procedure improves accuracy over zero-shot by 1.22 percentage points on A-OKVQA and 0.53 on AI2D, while breaking fewer correct answers than the retrievers. These findings suggest the potential value of studying multimodal ICL through the characteristics of queries themselves.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.