Visibility Is Not Authority: Visual Task Instruction Delegation in LVLMs
Abstract
Large vision–language models (LVLMs) can follow instructions embedded in images without user authorization, allowing visual content to redirect the user’s task. Such overreach can occur even when the instructions are harmless. However, multimodal safety research focuses primarily on harmfulness, leaving the authorization boundary underexplored. We formalize this challenge as visual task instruction delegation (VTID), distinguishing unauthorized non-execution, authorized-safe completion, and authorized-unsafe refusal. To assess this boundary, we introduce VTID-BENCH, a 30,000-example benchmark and alignment resource with controlled counterfactuals that separate user authorization from task safety. Evaluations reveal that natural and detailed-policy prompts fail to establish a reliable authorization boundary across four LVLMs. To investigate whether this boundary can be learned through alignment, we develop VTID-SFT, which explicitly separates authorization from safety to guide how models respond to instructions embedded in images. Across three LVLMs, VTID-SFT achieves over 95% three-state accuracy and improves authorized-safe task completion, raising LLaVA’s three-state accuracy from 26.33% to 95.30%. External evaluations show improved LLaVA safety while retaining at least 95.9% of its base performance on three general vision–language benchmarks. These findings establish user authorization as a distinct dimension of multimodal alignment and demonstrate that models can learn to respect it while retaining general capability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.