Hyperbolic Detoxification for Vision-Language Model Safety via Intent-Risk Calibration
Abstract
Improving the safety of vision–language models (VLMs) without compromising their utility on benign image–text inputs remains a fundamental challenge. A direct strategy is to first generate a draft response and then determine whether it should be rewritten based on its content. Although simple and effective, this response-first approach can be unreliable for ambiguous, gray-zone cases in which output-level signals alone are insufficient for robust safety decisions. To tackle the issue, we introduce Intent-Calibrated Hyperbolic Detoxification, a response-first safety framework that jointly assesses output safety and multimodal input risk. Our method first evaluates each draft using a response-level reward model: high-risk drafts are forcibly rewritten, whereas clearly safe drafts are returned unchanged. For drafts near the safety decision boundary, a hyperbolic encoder captures the hierarchical risk structure of the image–text input, and a learned multilayer-perceptron gate combines the resulting input-risk score with the response-level reward to determine whether rewriting is necessary. This design restricts input-driven intervention to genuinely ambiguous cases, thereby reducing unnecessary refusals and preserving benign responses. We evaluate the proposed method using attack success rate (ASR) and helpfulness rate (Helpful) on VLM safety benchmarks, together with accuracy on RealWorldQA. We additionally benchmark the underlying backbones and adapted safety baselines on MMMU and MMStar. Across the safety and RealWorldQA evaluations, we achieve a more favorable safety–utility trade-off than existing strong baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.