D-BISR: Multi-Prototype Diffusion Residual Scoring for Target-Data-Free Harmful Prompt Detection
Abstract
Large Language Models deployed in open-ended systems require input-side harmful prompt detection to block safety risks before generation. In practice, harmful prompts in target applications often differ from training data, and target-domain annotations are difficult to obtain. Detectors should therefore generalize to unseen domains without target-domain samples. While encoder-based models offer efficient input filtering, they suffer from local risk dilution, risk-pattern entanglement, and classification scores that fail to support stable low-FPR detection. To address these challenges, we propose D-BISR, an encoder-based framework for target-data-free harmful prompt detection. D-BISR uses token-level prototype attention to aggregate local harmful evidence and introduces multiple learnable risk prototypes to separate distinct risk patterns. Instead of relying on standard classification logits, it performs diffusion residual scoring to map risk evidence into a continuous, thresholdable safety metric. We evaluate D-BISR on ViSU, MMA, and SneakyPrompt under strict target-data-free settings. When trained solely on source-domain data and tested directly on unseen target domains, D-BISR achieves strong cross-dataset generalization and maintains stable low-FPR recall. On standard in-domain benchmarks, D-BISR outperforms existing encoder-based classifiers and anomaly detection baselines. Compared with larger safety guard models, D-BISR uses fewer parameters, less GPU memory, and achieves lower inference latency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.