acceptodds
Under review as a conference paper at ICLR 2027

Blackwell-Guided Logit Search (BLS): Optimizing Multi-Turn Jailbreaks with Subliminal Steering

Abstract

Automated multi-turn jailbreaks typically steer search with a single scalar judge score computed from target-generated responses. This evaluation loop is expensive and can become uninformative when safety-trained models produce near-identical refusals. We introduce Blackwell-Guided Logit Search (BLS), which instead probes the target model’s internal state relative to an attacker objective using next-token logit margins on a small set of indicative multiple-choice questions. This yields a vector-valued progress signal without sampling full target responses or training an auxiliary judge. BLS then applies Blackwell approachability to optimize attacker prompts sentence by sentence, adaptively emphasizing whichever probe dimension(s) are most behind. The resulting intermediate signal exposes "subliminal” steering effects of token sequences that may be invisible from surface text or to an external LLM judge. Against a safety-aligned 72B model behind a strict refusal system prompt at a matched attacker-token budget, vector BLS achieves substantially higher attack success rates across multiple scenarios (e.g., 37.3% vs. 7.1% for TAP and 0.0% for PAIR on one scenario) while requiring – fewer target-generated replies. BLS also outperforms TAP against a second model family hardened by representation rerouting, on a scenario adapted from StrongREJECT. We also study the limits of inference-time subliminal steering in isolation, showing that surprisingly, BLS can induce inconsistencies or persona shifts by chaining sentences sampled from a standard pretraining corpus whose surface content bears no apparent relation to the induced behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.