Safe Representation for Combined Image-Text Generation in Multi-Turn Conversations
Abstract
Multi-turn conversational interaction introduces complex security risks as harmful intentions become dispersed and concealed across dialogue steps. In unified models, this problem is exacerbated by the tight integration of image and text modalities, leading to safety misalignment during continuous generation. While current research filters explicit risks within single modalities, it overlooks the emergent harm induced by cross-modal interactions. We identify a critical risk pattern termed semantic mis-grounding, where standalone images and text appear benign, yet their combination conveys inappropriate thematic meanings, and the core capability at stake is whether an image is permitted to be interpreted under a specific thematic perspective. Empirical evaluations across five unified models confirm the prevalence of this issue. To address it, we propose Uni-SafeR, a framework for Unified Safe Representation that focuses on the associative semantics between modalities. Uni-SafeR performs safety alignment over full multi-turn interaction trajectories: the objective spans both the image generated in the first turn and the safety-aware textual response in the second turn, jointly shaping the contrast between safe and harmful responses and the contrast between benign and mis-grounding image generations. Training is instantiated on a dataset organized around four task-specific strategies, including diversified templates, benign anchors, counterfactual pairs, and multi-source images, which encourages the model to reason about the safety of guided multimodal alignments and to reject unjustified associations. Extensive experiments demonstrate that Uni-SafeR effectively reduces multimodal harm across various models while preserving, and even improving, their core generation and understanding capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.