ICL-R: Reconstructing the Computational Effects of Demonstrations for Implicit Multimodal In-Context Learning
Abstract
Multimodal in-context learning (ICL) enables multimodal large language models (MLLMs) to adapt to downstream tasks through interleaved image and text demonstrations. However, retaining these demonstrations not only increases the sequence length and latency at inference time but also makes predictions sensitive to how demonstrations are selected and presented. Although existing implicit ICL methods approximate demonstration effects through representation shifts or attention routing, they still suffer from coarse or partial reconstruction of these effects. To address these limitations, we first decompose the change in attention output caused by demonstrations into the attention routing shift, the value shift, their interaction, and information contributed directly by demonstration tokens. Based on this decomposition, we propose In-Context Learning via Reconstruction (ICL-R), a lightweight framework for jointly reconstructing these computational effects. Aiming at incomplete reconstruction under demonstration removal, we propose mass-restoring joint attention-value calibration which jointly calibrates attention routing and value representations while rescaling the attention output. To further extend supervision beyond isolated state matching, we introduce a Riemannian state transition objective that aligns both incoming states and the transitions induced by attention. Extensive experiments across multiple MLLMs and benchmarks show that ICL-R improves performance while maintaining an inference cost close to that of standard zero-shot inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.