acceptodds
Under review as a conference paper at ICLR 2027

JOLT: Joint Text–Image Search for Grounding Disruption

Abstract

Grounded multimodal reasoning systems select image regions as evidence for answering, exposing an attack interface when those selections are observable. We study a reference-assisted, query-bounded attack that jointly modifies questions and applies bounded image perturbations, using only returned boxes as target feedback. The challenge is that localization depends on the complete image–question pair: textual and visual changes must be evaluated together within a limited query budget. We introduce JOLT, which represents question variants and accumulated image perturbations as persistent paired states. Geometric feedback guides their joint search, allowing candidates to retain earlier image changes while exploring new input combinations without model internals or downstream answer feedback. We assess both localization and answer outcomes. Preliminary reported results across six datasets show that JOLT reduces mean answer accuracy from approximately 70.2% before attack to 29.0% after attack. These findings demonstrate downstream degradation in the evaluated systems and identify a robustness concern for multimodal pipelines that expose intermediate visual grounding and use selected regions as evidence for answering.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.