JailNewsGuard: Release-Time Risk Gating for Jailbreak-Induced Fake News
Abstract
While existing jailbreak defenses primarily target harmful instruction following or toxic content generation, relatively little attention has been paid to release-time misinformation containment. This distinction is particularly important because misinformation generation depends not only on refusal behavior but also on factual consistency, narrative framing, and semantic manipulation. We propose a three-stage defense system called JailNewsGuard that screens explicit templates and context-length anomalies before generation, then scores generated text for eight harmfulness dimensions: faithfulness, verifiability, malicious-instruction adherence, scope, scale, formality, subjectivity, and agitativeness. The defense operates strictly as an external containment gate without modifying model weights or improving internal safety alignment. Moreover, we provide theoretical support for false release bound, dominance of multi-dimensional vector gating, and generalization bound for hidden-state safety probe. We conduct experiments on JailNewsBench with 10000 news items, 7 attack conditions, 23 languages, and 34 regions, yielding 70000 paired baseline and defense records. Our method successfully suppresses the attack success rate from 68.29% to 4.39% for DeepSeek-V4-Pro, and from 17.40% to 1.09% for GPT-5.5, corresponding to relative reductions of 93.57% and 93.75%, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.