acceptodds
Under review as a conference paper at ICLR 2027

Direction-Indexed Evolutionary Red-Teaming of Large Language Models

Abstract

We introduce Direction-Indexed Evolutionary Red-Teaming (DIER), a black-box jailbreak method that learns its search coordinate from a target language model's own responses rather than from the text of the attack. A search that accumulates reusable framings, such as roles, pretexts and personas, has to decide which are close enough to recombine, and text does not carry that decision: across 21 target models, how a framing is worded says little about what it does, and a whole response is dominated by what was asked rather than by how it was asked. After each attempt DIER takes the displacement of the response from the request in a fixed sentence-embedding space and projects out the direction every response to that request shares; what survives is a signature of the tactic alone. That signature is cheap to measure and not predictable from a framing's text, so it can be read only after an attempt and used only to rank the library the search has already built. That ranking is the coordinate's only causal path into an attack: the direction never enters a prompt, is never decoded into words, and never touches the target's activations. On the 400-behavior HarmBench text benchmark at 32 target queries per behavior, DIER attacks at the state of the art: a mean attack success rate of 86.4 across the 16 targets with published baselines against 55.2 for the strongest attack in the HarmBench suite, level on the mean with the strongest prior strategy-library agent and ahead of it on all five targets where that agent stays below 90 percent. Replacing the ranking with a uniform random draw from the same library, with the attacker model, the library and the budget untouched, costs 13.5 to 39.0 points, placing the effect in the measured coordinate rather than in the language model that writes the attacks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.