MCPA: Modal Conflict Prompt Attack on VLMs for Autonomous Driving
Abstract
Vision-language models (VLMs) are increasingly used in autonomous driving (AD) systems for scene understanding and driving guidance. However, when user queries contain incorrect descriptions of objects, attributes, or relations, AD-VLMs may accept these false premises without verifying them against visual evidence, leading to unsafe driving suggestions. Existing datasets provide limited support for studying such vulnerabilities. Therefore, we construct DriveConflict, a dataset derived from DriveLM that introduces object, attribute, and relation conflicts between textual queries and visual scenes. Evaluations on general and AD-VLMs reveal a cross-modal premise verification deficiency, where models often prioritize answering queries over validating their textual premises. Motivated by this finding, we propose the Modal Conflict Prompt Attack (MCPA) for AD-VLMs in a black-box setting. Given a visual input and a user query, MCPA first identifies semantic anchors involving objects, attributes, or relations and selects replacements that preserve the semantic type but are unsupported by the visual evidence. It then constructs false premises through minimal semantic modification and embeds them into prompts designed to induce shifts in scene understanding or unsafe driving guidance. The resulting prompts remain linguistically natural while contradicting the visual scene. Experimental results demonstrate that MCPA achieves strong attack performance across datasets, models, and attack types. On NuScenes-QA, MCPA achieves an FPAR of 98.05%. On DriveLM, it achieves an FPAR of 99.80%, a DSR of 93.08%, and a UGR of 98.17%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.