TrajBridge: Trajectory-Guided Safety Recovery For Multimodal Language Models
Abstract
Multimodal large language models (MLLMs) extend language models with visual understanding, but the visual inputs can also disrupt safety behavior that is otherwise preserved under text-only input. In particular, the same MLLM may refuse a harmful request in text form yet comply once an image is introduced. We refer this phenomenon as vision-induced safety degradation (VD). Existing methods typically apply representation-level corrections through fixed intervention schemes. However, such fixed schemes may fail to adapt as the response evolves, causing correction to act at inappropriate locations or persist for an unsuitable duration. To better adapt correction to the evolving response, we propose TrajBridge, an inference-time framework for trajectory-guided safety recovery based on state-dependent feedback. TrajBridge first constructs layer-wise correction references offline to support both localization and correction. During online correction, it repeatedly reassesses the evolving response state to localize where intervention is most useful and applies bounded, state-dependent correction with adaptive strength. A separate correction controller governs when intervention should remain active or be released, preventing correction from being applied longer than necessary. By continuously coupling intervention decisions with the evolving response trajectory, TrajBridge turns static representation calibration into trajectory-guided, closed-loop safety recovery. Across paired VD evaluations, TrajBridge consistently improves safety recovery while preserving stable-safe behavior and general multimodal capabilities. Further analyses validate the importance of dynamic intervention location and duration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.