Toward Robust Multimodal Retrieval for MLLMs with Transport Alignment
Abstract
Multimodal retrieval ranks image-text candidates according to semantic relevance and supports applications such as image-to-text retrieval, text-to-image search, and composed retrieval. However, existing multimodal retrieval models are highly sensitive to adversarial perturbations on image queries. Even when the perturbation is almost invisible, it can move the query representation away from the original cross-modal semantic structure, distort matching between the query and candidates, and sharply degrade ranking performance. To address this problem, we propose ROTA, a training-free framework for robust multimodal retrieval with transport alignment. ROTA constructs transformed views of the attacked query and uses retrieval entropy to select reliable views, rather than relying on a single corrupted query representation. It further enriches candidate texts with descriptive views and aligns image-view and text-view distributions using optimal transport. The resulting retrieval score combines reconstructed query evidence with distribution-level image–text alignment, reducing reliance on brittle single-view matching. Furthermore, we analyze how perturbations in the OT cost matrix affect transport matching, which clarifies why distribution-level alignment can better tolerate small representation changes. Experiments across retrieval benchmarks, adversarial attacks, and retrieval backbones show that ROTA consistently improves performance over direct retrieval with attacked queries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.