acceptodds
Under review as a conference paper at ICLR 2027

SimJIR: Jailbreaking Aligned LLMs with Simple In-Context Rewards

Abstract

Large Language Model (LLM) safety is becoming increasingly important as LLMs are deployed in a growing number of real-world applications. Jailbreak attacks help expose vulnerabilities and foster the development of better alignment methods. Existing black-box attacks are costly: Many-Shot Jailbreaks require curated demonstrations, while Multi-Turn Jailbreaks rely on capable external attacker LLMs. We observe that a refusal emitted by a target LLM is a *free and on-policy* negative demonstration underutilized by existing attacks. Combining this observation with the recent finding that LLMs can improve their responses based on in-context rewards given to their previous attempts, we propose SimJIR (Simple Jailbreaking with In-context Rewards), a simple, multi-turn black-box attack that does not need any curated demonstrations or attacker LLMs. SimJIR assigns a negative reward to each past refusal and maintains a cumulative score across turns, penalizing the model until it decides to maximize its reward by producing a harmful response. We demonstrate empirically that SimJIR achieves competitive attack success at a fraction of the cost, can be adapted to bypass defenses through small edits, and can be used as a vision-language jailbreak out of the box.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.