acceptodds
Under review as a conference paper at ICLR 2027

JailNewsGuard: Release-Time Risk Gating for Jailbreak-Induced Fake News

Abstract

While existing jailbreak defenses primarily target harmful instruction following or toxic content generation, relatively little attention has been paid to release-time misinformation containment. This distinction is particularly important because misinformation generation depends not only on refusal behavior but also on factual consistency, narrative framing, and semantic manipulation. We propose a three-stage defense system called JailNewsGuard that screens explicit templates and context-length anomalies before generation, then scores generated text for eight harmfulness dimensions: faithfulness, verifiability, malicious-instruction adherence, scope, scale, formality, subjectivity, and agitativeness. The defense operates strictly as an external containment gate without modifying model weights or improving internal safety alignment. Moreover, we provide theoretical support for false release bound, dominance of multi-dimensional vector gating, and generalization bound for hidden-state safety probe. We conduct experiments on JailNewsBench with 10000 news items, 7 attack conditions, 23 languages, and 34 regions, yielding 70000 paired baseline and defense records. Our method successfully suppresses the attack success rate from 68.29% to 4.39% for DeepSeek-V4-Pro, and from 17.40% to 1.09% for GPT-5.5, corresponding to relative reductions of 93.57% and 93.75%, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.