Blackwell-Guided Logit Search (BLS): Optimizing Multi-Turn Jailbreaks with Subliminal Steering
Abstract
Automated multi-turn jailbreaks typically steer search with a single scalar judge score computed from target-generated responses. This evaluation loop is expensive and can become uninformative when safety-trained models produce near-identical refusals. We introduce Blackwell-Guided Logit Search (BLS), which instead probes the target model’s internal state relative to an attacker objective using next-token logit margins on a small set of indicative multiple-choice questions. This yields a vector-valued progress signal without sampling full target responses or training an auxiliary judge. BLS then applies Blackwell approachability to optimize attacker prompts sentence by sentence, adaptively emphasizing whichever probe dimension(s) are most behind. The resulting intermediate signal exposes "subliminal” steering effects of token sequences that may be invisible from surface text or to an external LLM judge. Against a safety-aligned 72B model behind a strict refusal system prompt at a matched attacker-token budget, vector BLS achieves substantially higher attack success rates across multiple scenarios (e.g., 37.3% vs. 7.1% for TAP and 0.0% for PAIR on one scenario) while requiring – fewer target-generated replies. BLS also outperforms TAP against a second model family hardened by representation rerouting, on a scenario adapted from StrongREJECT. We also study the limits of inference-time subliminal steering in isolation, showing that surprisingly, BLS can induce inconsistencies or persona shifts by chaining sentences sampled from a standard pretraining corpus whose surface content bears no apparent relation to the induced behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.