acceptodds
Under review as a conference paper at ICLR 2027

BlueAnchor: Reinforcing Omni-LLM Defenses Against Jailbreaking through Adaptive Semantic Anchoring

Abstract

Omni-modal large language models are increasingly used as interactive agents because they can jointly understand images, text, and audio. The same cross-modal reasoning capability creates a new jailbreak surface: harmful intent can be divided among individually benign-looking inputs and become apparent only after the model combines them. Existing defenses often inspect each modality independently, stack costly purification modules, or reject every flagged request, leaving cross-modal blind spots while reducing efficiency and benign utility. We present BlueAnchor, an input-level defense for frozen black-box Omni-LLMs. A lightweight Semantic Sentinel Agent examines the complete request and produces a structured action plan containing the risk category, affected modalities, intervention location, and strength. BlueAnchor executes this plan through localized visual anchors, short audio cues, and textual safety frames before a single target-model query. Reusable visual anchors are learned offline with a benign-centroid surrogate objective, while audio and text actions are selected from category-conditioned libraries. A conditional surrogate analysis connects routing quality, residual anchor distance, and surrogate-to-target transfer. Across visual, audio, and compositional jailbreak benchmarks, BlueAnchor consistently reduces attack success while preserving benign task utility. Controlled studies isolate the gains from routing and intervention, and additional evaluations demonstrate cross-model transfer, low inference overhead, and improved adaptive robustness through randomized anchor selection. redWarning: This paper contains potentially harmful content.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.