FOSTER: Fixed-Threshold Safety Topology Repair for VLM Jailbreak Detection
Abstract
Deployed jailbreak detectors for vision–language models (VLMs) must apply a single decision threshold, chosen before unseen attacks appear, to open-world traffic. We formalize this constraint as fixed-threshold transfer failure: an un- seen attack can remain rank-separable from benign inputs (high AUROC) while its scores drift below the pre-deployment threshold (low TPR). Despite this global displacement, unseen attacks preserve a transferable local unsafe topology in mid- to-late VLM hidden layers, so a limited number of trustworthy anchors can be sufficient to improve subsequent detection without retraining or threshold recal- ibration, while keeping the benign reference fixed. Building on this, we pro- pose FOSTER, which freezes the VLM, the safety evidence router, the benign reference, and all pre-deployment thresholds, while selectively expanding only a cross-verified unsafe topology bank through mode-preserving anchor admis- sion. Anchors are cross-checked, pseudo-labeled unsafe samples, admitted only when the current topology, the frozen initial topology, and a separately param- eterized safety estimator all satisfy their own pre-calibrated admission criteria, which are distinct from the prediction and routing thresholds. The three witnesses are complementary rather than statistically independent; their purpose is to re- duce self-reinforcing admission errors. We evaluate FOSTER under fixed cali- bration thresholds and streaming deployment on Qwen2.5-VL and LLaVA-v1.6. Under each detector’s own pre-deployment threshold, pooled attack recall rises from 10.5% to 80.7% on Qwen2.5-VL and from 6.8% to 71.3% on LLaVA-v1.6, with zero observed false positives on the 917-sample deployment benign pool (a 95% Wilson upper bound of 0.42%, not a guarantee of zero benign error). Struc- turally, freezing the admission witnesses fixes the eligibility set before deploy- ment, which yields a stream-independent outer envelope on any input’s online topology score—a conditional bound rather than a deployable certificate, since el- igibility does not imply admission. Because the benign reference bank is frozen by construction, benign-side drift is measured rather than bounded: under out-of- distribution benign pools and targeted benign contamination we report the result- ing change in false positives explicitly. Our code is available at an anonymous repository: https://anonymous.4open.science/r/FOSTER
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.