What Output Scores Miss: Auditing Joint Safety in Frozen Vision-Language Models
Abstract
Vision-language safety depends on interpreting images and text in context, as their combination can introduce risks absent from either modality alone. Recent benchmarks reveal that models can miss these compositional risks even when their unimodal judgments are correct. Evaluations based on final outputs identify such failures, but leave unresolved whether useful safety evidence remains recoverable from frozen internal states beyond the model’s output scores. We find that joint states and their contrasts with unimodal states provide additional label information, allowing some joint errors to be corrected without updating the backbone. Building on this observation, we introduce Joint Safety Calibration (JSC), a supervised offline framework for auditing this residual joint-safety evidence. JSC combines the same model’s output scores with frozen-state features to assess their incremental predictive value under a controlled comparison. Cross-model experiments on native three-class safety assessment and controlled analyses show improved joint prediction and more effective repair of joint errors. The resulting audit clarifies what information remains recoverable after joint prediction failures and identifies opportunities for improving joint safety judgments through frozen-state readouts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.