RefBind: Reference-Guided Object-Level Semantic Replacement in VLMs
Abstract
Targeted attacks on vision–language models can induce a target concept without fully replacing the source object. We study reference-guided source-to-target semantic replacement, in which an attacker perturbs a source image while limiting interference with unrelated scene content. Building on disentangled Value features, RefBind extracts source- and target-object semantic supports and combines textual redirection with reference-visual alignment. Textual supervision operates over a broader spatial scope, whereas reference-visual supervision is restricted to the source-object support. A task-compatible intermediate representation complements final-level visual alignment. Controlled spatial ablations show that this asymmetric allocation improves replacement relative to the tested alternatives. On COCO300, RefBind increases the macro-average judge-assessed full-success rate over V-Attack from 5.58% to 37.33% for captioning and from 8.11% to 38.28% for VQA under the single-surrogate setting. On the 150-case robotic-scene benchmark, the macro strict source-to-target replacement rate increases from 5.47% to 28.40%. These results characterize a transferable vulnerability in multimodal semantic interpretation and planning-related VLM outputs, while the effectiveness of such attacks in physical-world robotic execution remains to be validated.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.