Non-Intrusive Safety Alignment for Vision-Language Models via Automated Output Rewriting
Abstract
Vision-language models (VLMs) remain vulnerable to multimodal jailbreaks. The usual remedy, safety fine-tuning of the model's own parameters, is unavailable for black-box models and can degrade general capability; post-hoc classifiers avoid both problems but answer every violation with the same fixed refusal. We propose NIRA (Non-Intrusive Rewriting Alignment), which places a safety rewriter at the output terminal of a frozen target VLM. The rewriter reads the image, the query and the target's reply; it returns a safe reply unchanged and replaces a harmful one with a refusal written for that request. Its training data come from a fully automated pipeline. We probe the target with multimodal jailbreaks, judge the replies with a guard model, and expand every successful attack with a modality-aware, two-track augmentation engine that re-instantiates the attack's evasion principle as new text and as newly synthesized images. The target's own safe replies, kept as identity pairs, teach the rewriter when not to intervene; trained without them, it refuses nearly every benign query. On JailBreakV-28K against a 30B target, NIRA lowers the attack success rate from 33.2% to 0.7%, and its answer is judged no worse than the undefended one on 88.7% of benign queries. The same rewriter remains effective on four held-out attack benchmarks, and on all five attack sets its only failures are misses: it never turns a reply the target already answered safely into a harmful one, and none of its interventions on a harmful reply leaves it harmful. Applied unchanged to a second target from a different model family, the pipeline yields a rewriter with the same properties.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.