Visuals Matter: Benchmarking and Enhancing Multi-Modal Knowledge Editing under Visual Distractions
Abstract
Multimodal knowledge editing (MMKE) has emerged as a critical approach for updating outdated knowledge in Vision-Language Models (VLMs). For MMKE, the knowledge must be grounded in visual cues that identify the target subject, so that the knowledge can transfer to any image of that subject while unrelated knowledge is preserved. However, real-world images often contain irrelevant visual elements, which we term as visual distractions. Existing MMKE benchmarks do not systematically examine their effects, leaving an open question: how do visual distractions affect MMKE? We introduce DistractEdit, a controlled, semi-synthetic benchmark that evaluates the same knowledge updates across increasingly complex visual scenarios: Isolated Entity, Contextual Background, and Multi-Subject Interference. By preserving the target subject’s visual details under different scenarios, DistractEdit enables controlled comparisons that isolate the effect of distractions on editing performance. Our evaluation across MMKE methods and VLMs reveals that visual distractions undermine editing performance, particularly the generalization of edited knowledge to unseen images, and the degradation degree varies across methods, models, and distraction types. We further propose SACO, an agent-based method that refines visual inputs by mitigating distractions before knowledge editing, following a See, Analyze, Clean, and Optimize strategy. Our results demonstrate that mitigating visual distractions improves the performance, highlighting visual refinement as a promising direction. We present SACO as a starting point and anticipate that DistractEdit and SACO will facilitate future research on MMKE under visual distractions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.