acceptodds
Under review as a conference paper at ICLR 2027

Next Demonstration Prediction for Visual In-Context Learning

Abstract

Selecting visual in-context demonstrations requires choosing both examples and their order. Visual nearest-neighbor retrieval can produce redundant prompts, while diversity heuristics cannot adapt their selection to a downstream vision-language model (VLM). We propose Next Demonstration Prediction (NDP), an autoregressive selector initialized by behavioral cloning of maximal marginal relevance and then trained with PPO using a frozen VLM's task reward. On 14 benchmarks with Gemma 3 4B-it at , NDP achieves the lowest mean regression MAE and highest mean VQA accuracy among the evaluated image-only selectors, but trails DPP on classification. Benefits vary by task and are limited at small shot counts. Reordering NDP's selected examples reduces performance in the evaluated in-domain settings, whereas cross-dataset policy transfer can underperform simple heuristics. NDP therefore offers a task-dependent, ordered alternative to static retrieval, at the cost of substantial VLM-based policy training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.