TRACE: TASK-AWARE ADAPTIVE SELF-EVOLVING AGENTIC JAILBREAKING
Abstract
The rise of LLM agents introduces a new threat by enabling planning, coding, and even end-to-end execution of expert-level attack workflows. However, this threat remains underexplored and underestimated since (i) Safety alignment prevents LLMs from directly generating harmful instructions. (ii) Although existing jailbreak methods can bypass safety alignment, most of them mainly focus on eliciting harmful instructions but overlook LLM refusal against executing these harmful instructions. (iii) Even if the early execution steps succeed, any intermediate failure or dependency error will interrupt the workflow and force a restart of the entire attack process. To address these limitations, we propose TRACE, an agentic jailbreaking framework that reduces attack difficulty, maintains execution dependencies, induces the agent to execute harmful instructions, and resumes from prior states rather than restarting. Specifically, we first design around 20 procedural decomposition schemes that specify key operations and dependencies for different types of tasks, which can decompose a complex and harmful expert-level task into more executable and benign-looking subtask sequences. We measure these candidate sequences along two dimensions, i.e., harmfulness and difficulty, and select the optimal one for instantiating the attack. During the attack process, most of the subtasks in the selected sequence are benign and thus can be executed by the agent smoothly. For the remaining refused subtasks, we construct different subtask profiles and establish an evolving execution-oriented multi-turn jailbreak strategy library corresponding to these profiles. This library continues to expand based on successful execution trajectories and recurring failure patterns, which already includes more than 30 strategies after evolution. Given the profile of a refused subtask, TRACE can retrieve an appropriate strategy that specifies the corresponding intermediate goals, environment states, and relevant tools to gradually establish a plausible execution context through multi-turn interaction, thereby inducing agent execution. When a step fails, TRACE localizes the failed subtask and adapts the interaction strategy to resume from the failed subtask rather than restart from scratch. Extensive evaluations on AdvCUA show that TRACE outperforms existing jailbreak methods across multiple advanced LLM agents under both regular and defensive settings, improving attack success rate by more than 100% over the recent agentic jailbreaking baselines. Our code is available herehttps://anonymous.4open.science/r/TRACE_ICLR-36E5.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.