acceptodds
Under review as a conference paper at ICLR 2027

Synthetic Multimodal In-Context Learning

Abstract

One of the central questions in multimodal in-context learning concerns how to construct effective multimodal contexts that provide appropriate task and reasoning guidance for multimodal large language models (MLLMs). Existing methods largely build such contexts by selecting demonstrations from a predefined corpus, limiting their ability to provide task- and reasoning-aligned context when suitable examples are unavailable. To overcome this limitation, we introduce synthetic multimodal in-context learning, a paradigm that shifts multimodal context construction from fixed-corpus demonstration selection toward query-adaptive demonstration synthesis at inference time. To instantiate it, we propose SynMIC, a self-evolving agentic framework that synthesizes demonstrations aligned with the task and reasoning schema required for in-context inference on each target query. Specifically, it abstracts the target query into a synthesis specification capturing its content, task, and reasoning schemas, and uses the specification to plan diverse demonstrations by specifying new visual scenarios, queries, pre-committed answers, and answer-supporting visual constraints. These plans are realized through tool-mediated synthesis by adaptively assembling and executing toolchains over code-based rendering, model-based image generation/editing, and auxiliary perception. Meanwhile, a verification-guided self-evolution loop verifies synthesized demonstrations, resynthesizes invalid cases, and distills failure experience into an evolving skill that assists subsequent synthesis. Experiments across diverse benchmarks and MLLMs demonstrate consistent improvements over existing baselines, highlighting the promise of shifting multimodal context construction from corpus-bounded retrieval toward adaptive synthesis. Code will be publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.