ILLUMINATING THE PRESSURES THAT BREAK LLM AGENTS: EVOLUTIONARY SEARCH FOR ALIGNMENT THAT GENERALISES
Abstract
When finishing a task means breaking a rule, language-model agents sometimes break it, and how often depends heavily on the system prompt, which operators write to get work done rather than to state every rule. We hold a decision fixed (its facts, actions and correct answer) and search over the system prompt around it, the framing, with a quality-diversity algorithm over two interpretable features, the safeguard a framing relies on and the pressure it applies, maximising either how often the agent violates or how undecided it is (the entropy of its choice). The best framing found per scenario raises the violation rate of three open-weight models from about 0% to 68–99.5% on average, and instructions to violate and changed company goals are associated with the most violation. The entropy objective finds framings on which each model is split between the two actions. Because the answer stays fixed while the framing varies, training on many framings should push an agent to decide by what the action does rather than by shortcuts in the prompt. Reinforcement learning on the framings the search finds, instead of on a neutral prompt, cuts Qwen3-8B's violations under new framings of held-out scenarios from 15% to 5%, also for pressures kept out of training, and gpt-oss-20b's from 7.5% to 4.4%. Running the search again against the trained models finds no prompt that makes either model violate even half the time (1.5–4.5%, against 37–51% for models trained on the neutral prompt). On insider trading and Agentic Misalignment, Qwen3-8B trained this way violates less than an untrained model told to act ethically, and on insider trading also less than after training on the neutral prompt (2.1% vs. 10.2%). Sixteen different prompts per scenario train better than one prompt repeated, and evolving each prompt beyond a single edit lowers harm on AgentHarm further and is slightly ahead on most other measures. In an exploratory study, changing the training examples in which acting is right (half of each training set) traded harm against over-caution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.