acceptodds
Under review as a conference paper at ICLR 2027

Learning Where to Draw the Line: Boundary-Aware Memory for Multimodal Safety Alignment

Abstract

Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, making safety alignment a fundamental requirement. A key challenge is that multimodal safety judgments often depend on fine-grained interactions between visual and textual inputs, where highly similar queries may require different judgments due to subtle cross-modal differences. We refer to this fine-grained distinction as the safety boundary. Existing approaches still fall short in handling such boundary cases. Training-based defenses require large amounts of labeled data, while training-free methods often rely on static instructions or case-level self-reflection. Since they do not explicitly represent or validate the safety boundary, these methods are less effective and generalize poorly to new boundary cases. To address this challenge, we propose BoundSafe, a training-free framework that learns external boundary-aware memory from limited policy-labeled examples. BoundSafe operates in three phases: Contrastive Memory Distillation constructs boundary-aware memory from paired examples, Cross-Validated Refinement validates and improves the memory, and Memory-Guided Inference applies the refined memory at test time. Extensive experiments on multiple multimodal safety benchmarks show that BoundSafe consistently outperforms all baselines across target models ranging from small open-source MLLMs to large proprietary ones. For example, on the MSSBench Embodied benchmark with InternVL3-8B, BoundSafe achieves 80.1% and 30.9% relative improvements in CCR and Avg over the strongest baseline, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.