acceptodds
Under review as a conference paper at ICLR 2027

Actions Speak Louder than Answers: Privileged-Context Self-Distillation for VLM Safety

Abstract

Vision-language models (VLMs) that refuse harmful text requests often comply when a benign-looking request is paired with an image that makes it harmful. On-policy self-distillation (OPSD) supervises every token of a model's own responses with a teacher that sees privileged context, usually a complete reference response, but it is unclear what this context should contain for safety. A response-level decomposition shows that naming the harm does not ensure refusal: four Qwen3-VL-4B baselines that we train for safety still fail to refuse in 67.0–82.6% of the responses that mention the harmful image content, close to the base model's 79.2%. We therefore give the teacher the action instead of an answer. Safety-Axis-Aware OPSD (SA²-OPSD) maps the image, request, and pair safety labels of each training pair to a short instruction that states whether to answer or refuse and where the risk lies, and the student, which sees only the image and request, learns from the teacher's next-token distributions on its own responses. Across two model scales and three safety benchmarks, SA²-OPSD achieves the lowest attack success rate (ASR), the share of unsafe inputs not clearly refused, in five of six settings. Against OPSD in the same pipeline, it lowers HoliSafe-Bench mean ASR over three runs from 37.7 to 25.8% at 4B and from 32.2 to 25.8% at 8B, and, in single 4B runs, ASR among responses that mention the harm from 68.6 to 42.6%. It also violates less often under a safe-completion rule and stays below the base model and OPSD on three further backbones, while refusing about 1% of fully benign inputs and keeping average capability within 0.6 points of the base models. OPSD remains 7.5 to 10.6 points behind at 4B even with references that agree with the pair labels or that the base model writes under the SA²-OPSD instruction, whereas an instruction that states only the action nearly matches SA²-OPSD: for safety self-distillation, stating the action works better than showing an answer. Code and data will be released. Warning: this paper may include examples of harmful content.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.