Distilling Reasoning Chain and Chain-Local Experiences for Few-Shot Multimodal In-Context Learning
Abstract
Multimodal in-context learning (ICL) offers a promising avenue for adapting Large Vision-Language Models (LVLMs) to specialized downstream vision-language tasks with only a few in-context demonstrations (ICDs). However, existing few-shot multimodal ICL methods either rely on superficial image-question-answer imitation or induce ICD-individual reasoning guidance, limiting their ability to capture shared task-level reasoning structures. Moreover, these methods are largely inherited from pure language scenarios, overlooking perceptual bottlenecks of models in multimodal tasks and failing to provide sufficient perceptual support for reliable step-wise execution. In this paper, we propose Distill-ICL, a training-free framework that distills few-shot ICDs into two complementary forms of guidance: a Normative Reasoning Chain and Chain-Local Experiences. Specifically, we introduce ICDs-MCTS, an ICD-set-oriented variant of Monte Carlo Tree Search that collectively expands and evaluates reasoning paths over the entire demonstration set, inducing a shared reasoning chain that specifies the sequence of reasoning steps for the target task. To support accurate visual perception when executing these steps, Distill-ICL further performs controlled rollouts along the induced chain and distills local guidance at critical nodes: Inter-ICD Pivotal Experiences capture answer-discriminative visual attributes to guide critical decisions, while Intra-ICD Fragile Experiences reinforce visual attributes prone to misinterpretation to help mitigate step-local perceptual errors. Extensive experiments on five specialized multimodal benchmarks show that Distill-ICL consistently outperforms existing training-free ICL methods across various LVLMs, while exhibiting favorable computational efficiency and improved robustness to ICD selection and ordering perturbations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.