On Evaluating the Adversarial Robustness of Foundation Models for Multimodal Entity Linking
Abstract
Multimodal entity linking (MEL) aims to identify a target entity from a set of semantically related candidates using visual and textual evidence. Unlike conventional image classification, MEL requires fine-grained candidate disambiguation, and small changes in cross-modal similarity may reverse the ranking between the correct entity and hard negatives. However, the adversarial robustness of MEL remains underexplored. We conduct a systematic evaluation of visual adversarial robustness under two candidate-constrained settings: Image-to-Text Entity Linking (I2T-EL) and Image+Text-to-Text Entity Linking (IT2T-EL). Our evaluation covers general-purpose vision-language encoders, multimodal large language models, three adversarial attacks, and five MEL benchmarks. The results show that visual perturbations can substantially reduce linking accuracy, although the degradation varies across models, datasets, and attack strengths. We further observe that informative textual context can partially preserve entity-level semantics under visual attacks. Motivated by this observation, we propose LLM-RetLink, a training-free framework that converts visual evidence into multiple discrete entity hypotheses, retrieves candidate-grounded knowledge, filters the retrieved evidence, and performs evidence-aware entity linking. Experiments show that LLM-RetLink improves both clean and adversarial linking accuracy by 0.4%-35.7% across the evaluated settings. Our study provides an adversarial MEL benchmark and highlights a new direction for improving entity-level robustness through discrete semantic grounding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.