acceptodds
Under review as a conference paper at ICLR 2027

CICD-Guard: Contextual Intent Consensus for Robust Defense against LLM Jailbreak Attacks

Abstract

Large language models (LLMs) are increasingly deployed in safety-critical applications, yet jailbreak attacks can conceal harmful intent through contextual packaging such as role-playing, fictional scenarios, or seemingly benign motivations. Existing input-level defenses often struggle to balance Defense Success Rate (DSR) and Over-refusal Rate (ORR), either missing disguised attacks or excessively refusing legitimate requests. We propose CICD-Guard, a parameter-free, black-box defense framework that exposes contextual intent through semantic rewriting and performs consensus-based risk assessment. Given an input prompt, CICD-Guard generates five semantically related rewrites that preserve the underlying objective while removing unnecessary contextual packaging. The original prompt and rewrites are then jointly evaluated by multiple independent LLM-based safety scorers using a continuous risk scale. A decision module combines absolute risk thresholds with relative risk changes between the original and rewritten prompts to adaptively intercept the request, preserve the original prompt, or return a more explicit representation with a safety warning. Our experiments show that CICD-Guard achieves average DSRs of 91.44% and 100% on the two target models, with ORRs of only 8.00% and 12.00%, respectively. Compared with the baselines, CICD-Guard provides a substantially more favorable DSR–ORR trade-off: on Llama-3-8B-Uncensored, it improves average DSR by 30.30 percentage points while reducing ORR by 10.04 percentage points; on DeepSeek-V4-Flash, it further improves average DSR by 15.67 percentage points while reducing ORR by 10.22 percentage points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.