AI Attacks AI: Agent Self-Jailbreaking through Malicious Self-Improvement
Abstract
As AI systems move from generating content to executing tasks, safety research has expanded from harmful model responses to harmful agent actions. Meanwhile, advances in self-improvement allow agents to acquire knowledge and modify their own implementations to improve performance. Since these modifications can also affect safety behavior, they raise the question of whether an agent can learn to change its response to a request it initially refused. We study agent self-jailbreaking, in which an agent acquires knowledge and modifies its own system to comply with a harmful request it initially refused. To examine how existing agents carry out this process, we introduce , a research harness that combines web search and local exploration with iterative self-modification. Task-wise research uses feedback from retrying the original request to guide further modifications. The modified agents are then evaluated on unseen tasks to assess transfer. Across five agent frameworks and three open-weight models, task-wise self-jailbreaking raises compliance from 0.9% to 70.9%. On unseen requests, the modified agents collectively achieve 97.0% compliance, compared with 1.8% for the original agents. Autonomous research on individual harmful requests can therefore change an agent's safety behavior across a broader range of tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.