acceptodds
Under review as a conference paper at ICLR 2027

CARS: Convergence-Aware Representation Steering for Multimodal Large Language Model Safety

Abstract

Multimodal Large Language Models (MLLMs) have achieved remarkable vision-language understanding through extensive pre-training, yet remain critically vulnerable to harmful visual inputs. Existing defenses either require costly retraining or introduce substantial overhead. To address this critical vulnerability, we introduce Convergence-Aware Representation Steering (CARS), an inference-time method exploiting natural cross-modal representation evolution during generation. Through systematic empirical analysis, we discover delayed refusal—a phenomenon where text-based safety steering often initially fails on harmful images but progressively activates as visual representations converge toward text-like patterns. CARS is designed to actively monitor this convergence and strategically enhances safety steering only when visual representations sufficiently align with text-based safety patterns. Experiments across FigStep, MM-SafetyBench, ToViLaG, and Jailbreak-28K show CARS achieves 92.3% average defense success rate—substantially outperforming state-of-the-art methods—while maintaining near-baseline utility with minimal degradation (0.1-1.3%) on ScienceQA and MM-Vet. Our work demonstrates that respecting natural representation dynamics enables superior safety-utility tradeoffs without architectural modifications or retraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.