Aligning with the Wrong Intent: How User-Intent Ambiguity Enables Semantic Attacks on Multimodal Agents
Abstract
As large language models evolve into autonomous agents, semantic attacks can redirect their actions away from the user's intended objective. Existing studies mainly analyze specific attack patterns, leaving unclear how incomplete user instructions affect vulnerability to these attacks. We study this problem from the perspective of intent alignment: malicious contextual information may be interpreted as evidence about what the user wants. We introduce AmbiSecBench, which systematically varies instruction completeness while holding environmental injections fixed across variants of the same decision state. Across GUI, Web, and embodied environments, we measure attack outcomes, candidate-based support for the original intent, and the relative preference for attacker-induced actions. Incomplete instructions consistently reduce support for the original intent and increase attacker-relative action preferences across the evaluated model–domain pairs. Paired clean and attacked comparisons further show that amplification of the attack-induced support loss depends on the domain. An intent-auditing prompt reduces attack success, although its effect on clean-task utility remains untested. These findings characterize how instruction completeness affects agent security and motivate defenses that distinguish user authorization from untrusted contextual information.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.