Know Your Target before Striking: A Profile-Aware Adaptive Jailbreaking Framework for LLMs
Abstract
Safety-aligned large language models (LLMs) remain vulnerable to jailbreak attacks. By exploiting an auxiliary LLM, Automatic jailbreak methods seek to generate adversarial prompts that induce LLMs to bypass their safety safeguards. However, most existing methods formulate jailbreaking as a target-model-agnostic prompt optimization problem. They therefore explore the search space without exploiting the behavioral characteristics of the target model. This unguided optimization often leads to unstable attack performance and requires a substantial query budget to identify successful jailbreak prompts. To address these limitations, we propose SafetySurface, a two-phase jailbreak framework that combines target-model behavioral profiling with profile-aware prompt optimization. In Phase 1, SafetySurface diagnoses the target model's safety behavior using a carefully designed probe dataset and constructs a behavioral profile that captures its response patterns. In Phase 2, an attacker model leverages this profile to steer prompt optimization toward more promising regions of the search space, enabling more targeted and efficient jailbreak generation. Experiments on AdvBench and HarmBench across multiple safety-aligned LLMs demonstrate that SafetySurface consistently improves attack success rate, query efficiency, and stability over strong baselines, achieving up to 100% attack success rate while substantially reducing query requirements.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.