acceptodds
Under review as a conference paper at ICLR 2027

The First Unsafe Token: When and Where Jailbreaks Take Over Language Models

Abstract

Aligned language models can refuse a harmful request yet comply when the same behavior is wrapped in a jailbreak. Recent work shows that refusal and jailbreak signals evolve across decoding trajectories, but a sharper causal question remains: when does a successful jailbreak cross a localized transition that controls subsequent generation? We introduce TIDE (Tokenwise Intervention for Decision Emergence), a layer by token analysis that replays an identical successful response prefix under a refused harmful prompt and its jailbreak, isolating prompt conditioned state divergence from lexical divergence. The First Unsafe Token (FUT) is localized from a development fixed, prefix aligned boundary score, and causal patching then verifies that final refusal and compliance is maximally sensitive near that boundary, including on natural free running trajectories. Our fixed protocol uses 160 behavior disjoint harmful requests, five attack families, four open weight model families, and 300 outer family test attempts per model. We observe an early boundary centered near token 4, with successful attacks crossing by token 6 far more often than failed attacks. This structure motivates TIDE-Guard, which reads the first four hidden states. Under nested held out attack family evaluation, TIDE-Guard reaches 0.91 AUROC, versus 0.80 for the strongest evaluated published baseline and 0.81 for a capacity matched four token state probe. At the endpoint, TIDE-Guard reduces jailbreak ASR from 64% to 27% at 4.2% benign refusal. Across four model families, the same early boundary structure reappears under model specific calibration.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.