JOLT: Joint Text–Image Search for Grounding Disruption
Abstract
Grounded multimodal reasoning systems select image regions as evidence for answering, exposing an attack interface when those selections are observable. We study a reference-assisted, query-bounded attack that jointly modifies questions and applies bounded image perturbations, using only returned boxes as target feedback. The challenge is that localization depends on the complete image–question pair: textual and visual changes must be evaluated together within a limited query budget. We introduce JOLT, which represents question variants and accumulated image perturbations as persistent paired states. Geometric feedback guides their joint search, allowing candidates to retain earlier image changes while exploring new input combinations without model internals or downstream answer feedback. We assess both localization and answer outcomes. Preliminary reported results across six datasets show that JOLT reduces mean answer accuracy from approximately 70.2% before attack to 29.0% after attack. These findings demonstrate downstream degradation in the evaluated systems and identify a robustness concern for multimodal pipelines that expose intermediate visual grounding and use selected regions as evidence for answering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.