acceptodds
Under review as a conference paper at ICLR 2027

ProbeFool: Evading LLM Jailbreak Activation Monitors with Probe-Agnostic Attention Routing

Abstract

Activation probes are increasingly proposed as lightweight monitors for harmful language-model requests. We ask whether attention at the monitored token is a usable proxy for the probe's own signal: whether the representation a monitor reads can be moved without ever optimizing against it. We introduce , a continuous-embedding attack that places a fixed benign distractor before an unchanged jailbreak prompt and optimizes a trainable steerer after it from frozen base-model signals alone, raising the monitored token's attention to the distractor while preserving the original response distribution. The probe is a decision-only black box: its decision would tell a deployed attacker when to stop, but it never enters the update. Against logistic-regression and MLP probes on Llama-3.2-3B-Instruct, Llama-3-8B-Instruct, and Mistral-7B-Instruct-v0.2, we conservatively report terminal-step results averaged over every transformer-layer probe. With the best of three distractor–steerer initializations selected on each model–probe track by joint success, the attack achieves 77–93% evasion with 52–77% joint success, meaning evasion and a still-harmful response on the same prompt, and it transfers across broad layer ranges and to unseen prompts. Controlled interventions locate the mechanism. Transplanting the attack's monitored-token attention onto an unoptimized wrapper recovers most of its evasion, while the optimized embedding under clean attention evades almost nothing. At fixed mass, which heads carry it decides the outcome.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.