acceptodds
Under review as a conference paper at ICLR 2027

Discovering jailbreaks via curriculum of models

Abstract

Automated jailbreak discovery aims to find prompts that elicit prohibited behaviors from safety-trained language models, helping developers identify weaknesses before deployment. One class of approaches trains an investigator model to query a target model, using reinforcement learning to reward prompts that successfully elicit a specified behavior. However, training directly against a safety-trained target is difficult because most prompts are refused, leaving the investigator with a sparse reward signal. Prior work addresses this problem with the propensity bound (PRBO) (Chowdhury et al., 2025), a proxy reward estimated by sampling responses from a modified target that exhibits the behavior more readily, and then adjusting for how likely those responses would have been under the original target. We propose Curriculum of Models (CoM), an alternative method that trains the investigator against a sequence of modified versions of the target model, from which the behavior is progressively harder to elicit. The curriculum begins with a version that rarely refuses harmful requests; once the investigator effectively elicits the behavior from the current version, training advances to the next, harder version, eventually reaching the original target model. We evaluate CoM on single- and multi-behavior elicitation tasks involving CBRN materials, explosives, and illicit drugs using two open-weight targets, gemma-4-31b-it and gpt-oss-120b, and find that CoM successfully discovers jailbreaks and achieves performance competitive with PRBO. Our results suggest that a simple curriculum over target models offers a promising alternative for addressing sparse rewards in rare-behavior elicitation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.