Visual Distraction Undermines Moral Reasoning in Vision-Language Models
Abstract
Moral reasoning is fundamental to safe Artificial Intelligence (AI), yet ensuring its consistency across modalities becomes critical as AI systems evolve from text-based assistants to embodied agents. Current safety techniques demonstrate success in textual contexts, but concerns remain about generalization to visual inputs. Existing moral evaluation benchmarks rely on text-only formats and lack systematic control over variables that influence moral decision-making. Here we show that visual inputs substantially alter moral decision-making in state-of-the-art (SOTA) Vision-Language Models (VLMs), revealing limitations in text-based safety alignment. We introduce Moral Dilemma Simulation (MDS), a multimodal benchmark grounded in Moral Foundation Theory (MFT) that enables mechanistic analysis through orthogonal manipulation of visual and contextual variables. The evaluation reveals that the vision modality induces more intuition-like decision patterns and weaker deliberative sensitivity than text-only contexts. These findings expose critical fragilities where language-tuned safety filters fail to constrain visual processing, demonstrating the urgent need for multimodal safety alignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.