acceptodds
Under review as a conference paper at ICLR 2027

Do Jailbreaks Leave Identifiable Behavioral Footprints? Detection and Localization through Attention Occupancy

Abstract

Jailbreak detectors typically render a single prompt-level decision, failing to pinpoint where adversarial manipulation occurs. We investigate whether a model's attention during early decoding can simultaneously support jailbreak detection and attack span localization without training auxiliary classifiers or updating model weights. Our approach isolates user-directed attention, filters attention sinks, and incorporates attention entropy to dynamically weight decoding queries according to their effective support size. Within individual heads, token profiles are robustly standardized using the median and median absolute deviation (MAD). A small set of task-relevant heads selected on development prompts provides token-level anomaly scores that support prompt-level detection and adversarial span localization. Operating on a frozen target model, our method separates the use of prompt labels for head selection from threshold calibration, which is performed on an independent non-attack split. Evaluated against GCG and HarmBench AutoDAN attacks across three target models, our detector identifies 87.00%–98.45% of attack attempts while flagging at most one of 17 benign references per target. It also yields the lowest judged attack success rate among the compared defenses and localizes annotated adversarial content in 86.50%–97.94% of evaluated attacks. These results show that early-decoding attention provides a useful, non-causal positional signal for jointly detecting and localizing jailbreak attacks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.