acceptodds
Under review as a conference paper at ICLR 2027

Improving Multimodal Many-Shot Learning with Needle-in-a-Haystack Training

Abstract

Vision-language models have shown significant promise in combining text and image modalities, yet their ability to effectively utilize multimodal many-shot context remains understudied. We study the performance of VLMs on multimodal many-shot classification problems and observe that their performance is inconsistent. On some datasets the classification performance improves linearly for every doubling of the number of shots, while for others, the performance outright degrades as the number of shots increases. Surprisingly, we find that finetuning these models on a set of multimodal many-shot classification tasks does not improve their performance on previously unseen many-shot classification tasks. Instead, we demonstrate that finetuning the models on a simpler but related needle-in-a-haystack problem, where the image to be classified and its label (needle) are hidden in the context (the haystack), significantly improves the performance and scaling with the number of shots. We analyze the attention patterns of these models to demonstrate evidence that needle-in-a-haystack training provides a strong learning signal to search the entire context, which changes the positional bias of the learned attention pattern and thus allows the model to generalize to many-shot settings. Overall, our results show that just 200 gradient steps of needle-in-a-haystack training may greatly improve performance on many-shot learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.