ACD-Retrieval: Towards Multimodal Knowledge Transfer via Aligned Cross-domain Demonstration Retrieval
Abstract
Multimodal large language models (MLLMs) have achieved remarkable progress in complex reasoning tasks, yet existing reasoning enhancement methods still rely on high-quality in-domain annotations or expert demonstrations. We observe that, even when in-domain demonstrations are unavailable, demonstrations from other domains can provide transferable reasoning knowledge for multimodal ICL. However, existing multimodal retrieval methods mainly rely on simple embedding similarity, which may not effectively reflect how well a cross-domain demonstration transfers to the target reasoning task, potentially leading to unstable retrieval. To address this challenge, we propose **ACD-Retrieval**, a unified framework for multimodal cross-domain ICL. ACD-Retrieval performs domain-wise centering and SVD-based subspace projection to obtain more reliable cross-domain similarity estimation, while a self-selection mechanism suppresses negative transfer during inference. Extensive experiments demonstrate that our method consistently achieves positive transfer across nine MLLMs, while exhibiting strong generalization across diverse multimodal cross-domain settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.