Compositional Unalignment: Auditing Safety in Large Vision-Language Models
Abstract
Large vision-language models (LVLMs) are vulnerable to compositional harm: an image and a prompt that are benign in isolation become harmful through the model’s cross-modal reasoning. Recent alignment methods report strong scores on compositional safety benchmarks, but assessing the outputs by an LLM judge only measures observed behavior and cannot reveal how robustly safety is integrated into the model. We argue that un-alignment provides a complementary audit of alignment robustness. We first show that minimal supervised fine-tuning achieves an LLM-judge score comparable to a multi-stage reinforcement-learning pipeline, showing that similar benchmark scores can arise from only a fraction of computations and cannot establish alignment robustness by themselves.We next evaluate whether existing representation-steering methods can serve as training-free unalignment mechanisms. We find that recognition and refusal of compositional harm are decoupled: nine standard direction-finding methods detect compositional harm with up to AUC = 0.99, yet ablating the recovered directions either leaves refusal intact or destroys model coherence, failing to properly un-align. We therefore propose CMaG, a training-free un-alignment method for compositional safety. CMaG obtains the refusal controlling direction causally: we measure refusal intention using a refusal-versus-compliance log-probability margin, compute its gradient with respect to mid-layer activations, and contrast them against partially benign pairs to isolate the cross-modal interaction term. A single backward pass over 80 examples yields a direction along which steering monotonically converts refusal into compliance. On Qwen2.5-VL-7B and LLaVA-1.5-7B, CMaG reduces compositional safety scores by up to 65 points while over 98% of outputs remain coherent, outperforming gradient-based baselines. We find that the recovered direction also transfers to text-explicit harmful inputs. CMaG is also capable of un-aligning models hardened by defensive training. These results suggest that compositional safety can appear behaviorally strong while remaining fragile to targeted internal intervention, and un-alignment is a necessary and complementary audit alongside benchmark scores. Warning: this paper contains unsafe model outputs, including references to self-harm.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.