acceptodds
Under review as a conference paper at ICLR 2027

For Research Purpose Only - Jailbreaking Large Language Models with Jailbreaking Research

Abstract

We introduce a simple yet powerful jailbreak framework for LLMs: Jailbreaking with Alignment Research Context (JARC). JARC interacts with the victim LLM through a two-stage process. The first stage establishes the context of research on LLM alignment and then requests for answering questions in a neutralized template. The second stage asks the victim model to fill in the template blanks, framing the task as a cloze test puzzle via blank-filling function calls. As a static attack strategy that does not rely on an auxiliary red-teaming model, JARC demonstrates high attack effectiveness with significantly lower token usage compared to existing attacks. Specifically, JARC achieves on average 80.7% Attack Success Rate (ASR) when attacking heavily safely-aligned frontier LLMs: GPT-5.1-Codex-Max, Gemini 3.1 Pro, and Claude Sonnet 4.5, yielding a 40%+ ASR improvement compared to state-of-the-art baseline attacks. We further encode the design pattern of JARC as a lightweight skill-based workflow and utilize Claude Code to instantiate the framework, demonstrating that agentic harnesses can autonomously generate diverse realizations of the attack while retaining strong effectiveness. Beyond attack effectiveness, we systematically investigate the robustness, contributing factors, and execution stability of JARC. Our study reveals the following two key messages: (1) *The scaling of agentic capabilities may not result in proportionally enhanced safety alignment.* Safety mechanisms that resist direct harmful requests may become less reliable when the same generation task is decomposed across stages, transferred across models or sessions, or expressed through agentic interfaces. As a result, agentic capabilities such as tool calling and multi-stage task execution, along with disparities in alignment efforts across LLMs can translate to new attack surfaces. (2) *Agentic harnesses can enable the "scaling" of attacks themselves.* These "dual-use" harnesses can be easily re-engineered to generate automated, diverse, highly effective jailbreaks at low cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.