acceptodds
Under review as a conference paper at ICLR 2027

Adaptive Demonstration Count for Multimodal In-Context Learning

Abstract

Multimodal in-context learning (MICL) enables multimodal large language models (MLLMs) to adapt to downstream tasks through visual-textual demonstrations without parameter updates. Existing MICL methods mainly focus on which demonstrations to retrieve, while typically using a fixed demonstration count for all queries. However, we observe that the optimal demonstration count varies substantially across query images, tasks, retrieval strategies, and target MLLMs. This problem is particularly important in multimodal settings, where redundant demonstrations introduce substantial visual-token overhead and increase inference latency. To address this issue, we propose Query-Adaptive Demonstration Count (QADC), a performance-driven and plug-and-play framework that dynamically determines the demonstration count for each multimodal query. We adopt a small–large model collaboration paradigm. Specifically, we additionally train a lightweight performance predictor to assist MLLMs in adaptively selecting the demonstration count. The performance predictor estimates the performance of candidate contexts constructed with different numbers of demonstrations and selects the demonstration count with the highest predicted performance. Extensive experiments on image classification, image captioning, and visual question answering benchmarks demonstrate that QADC consistently obtains the best or second-best performance across redX open-source MLLMs and redX commercial MLLMs. By selecting shorter contexts when additional demonstrations are unlikely to be beneficial, QADC reduces the average token cost by 60.3% and inference time by 62.1% compared with a strong retrieval baseline. It also exhibits strong generalization across MICL methods, datasets, prompt templates, and target MLLMs. QADC can serve as a plug-and-play module to improve the performance of other MICL methods. Moreover, QADC can be applied in many-shot demonstration settings and is also applicable when only a small amount of training data, such as 30 training examples, is available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.