acceptodds
Under review as a conference paper at ICLR 2027

Know Your Target before Striking: A Profile-Aware Adaptive Jailbreaking Framework for LLMs

Abstract

Safety-aligned large language models (LLMs) remain vulnerable to jailbreak attacks. By exploiting an auxiliary LLM, Automatic jailbreak methods seek to generate adversarial prompts that induce LLMs to bypass their safety safeguards. However, most existing methods formulate jailbreaking as a target-model-agnostic prompt optimization problem. They therefore explore the search space without exploiting the behavioral characteristics of the target model. This unguided optimization often leads to unstable attack performance and requires a substantial query budget to identify successful jailbreak prompts. To address these limitations, we propose SafetySurface, a two-phase jailbreak framework that combines target-model behavioral profiling with profile-aware prompt optimization. In Phase 1, SafetySurface diagnoses the target model's safety behavior using a carefully designed probe dataset and constructs a behavioral profile that captures its response patterns. In Phase 2, an attacker model leverages this profile to steer prompt optimization toward more promising regions of the search space, enabling more targeted and efficient jailbreak generation. Experiments on AdvBench and HarmBench across multiple safety-aligned LLMs demonstrate that SafetySurface consistently improves attack success rate, query efficiency, and stability over strong baselines, achieving up to 100% attack success rate while substantially reducing query requirements.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.