acceptodds
Under review as a conference paper at ICLR 2027

CARD: CATEGORY-ALIGNED REFUSAL DIRECTIONS FOR JAILBREAK ATTACKS

Abstract

Research on jailbreaking large language models (LLMs) has attracted increasing attention in recent years. However, existing jailbreak attacks are often evaluated on datasets with coarse-grained harm categorization, leaving their effectiveness across diverse harmful requests insufficiently understood. We observe that their attack success rate (ASR) drops substantially on safety datasets with fine-grained and diverse semantic categories. To address this challenge, we propose CARD, a jailbreak framework guided by category-aligned refusal directions. CARD extracts a refusal direction for each harm category from harmful–harmless prompt contrasts, selects the corresponding direction via a lightweight hidden-state classifier when the target category is unknown, and incorporates the selected direction into adversarial suffix optimization. Specifically, CARD employs a projection-guided suffix optimization strategy: the projection term penalizes hidden-state components aligned with the category-aligned refusal direction while retaining the target-likelihood objective. Experiments show that CARD outperforms most of the evaluated attack baselines in ASR on BeaverTails, a safety benchmark spanning harm categories. Moreover, CARD generalizes to AdvBench and HarmBench, achieving up to 79.25$% ASR, respectively, despite differences in their harm-category taxonomies and annotation schemes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.