Hidden Harm in Non-Atomic Demonstrations: Selective Counterfactual Component Intervention for Multimodal In-Context Learning
Abstract
Multimodal in-context learning (MICL) commonly constructs and evaluates in-context sequences at the level of complete image-text demonstrations, although prior studies show that visual and textual information can contribute differently across tasks. We ask a finer-grained question: for a query, do the image and text within the same demonstration exert the same effect? They need not. Their query-conditioned utilities can differ, and sometimes disagree in sign, such that a demonstration appearing useful at the whole-demonstration level may still contain a component that harms the prediction. We refer to this failure mode as hidden component harm, arising from demonstration non-atomicity. To mitigate it, we propose Selective Counterfactual Component Intervention (SCCI), a training-free framework that decomposes each demonstration into visual and textual components and derives query-specific counterfactual signals through leave-one-component-out inference. A reliability-aware decision rule intervenes only when the evidence is sufficiently informative, while otherwise preserving the context. Across six vision-language benchmarks and two primary multimodal language models, Qwen2.5-VL-7B and InternVL2.5-8B, our method consistently improves 4-shot MICL over the unmodified context and outperforms both random component removal and matched whole-demonstration removal. Further analyses of image-text utility disagreement, hidden component harm cases, and whole-demonstration attribution support the value of semantic component granularity, with gains remaining robust across demonstration sampling seeds and order permutations. Anonymous code is available at https://anonymous.4open.science/r/SCCI-3F45/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.