Amplify to Detect: Gradient-Guided Weight Interventions for Multimodal Jailbreak Detection
Abstract
Vision–language models (VLMs) can be jailbroken by malicious intent conveyed through text, images, or their interaction. Detecting such inputs before harmful generation is therefore important, but detectors must also avoid falsely flagging benign requests. Existing hidden-state detectors use VLM representations as fixed features, leaving unclear which internal computations produce harmful–benign separation and whether these signals can be strengthened. We introduce Amplify to Detect, a frozen-model intervention for multimodal jailbreak detection. Our method places multiplicative scalers on output channels of VLM linear projections and uses first-order sensitivity of a paired harmful–benign representation-separation objective to identify influential channels. We amplify channels predicted to increase separation, then apply hidden-state detectors to the resulting representations. Across three open-weight VLMs and multiple detectors, targeted amplification generally improves malicious-intent detection under a stringent 5% false-positive-rate constraint. For HiddenDetect, TPR at this operating point increases by 18.8–57.2 percentage points across the three models. Direction ablations further show that reversing the predicted intervention direction reverses its effect, while random would decrease the performance. These findings suggest that safety-relevant hidden-state signals arise partly from identifiable internal computations and can be selectively strengthened for more reliable jailbreak detection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.