J-Anchor: Defending Against Prefill Attacks via J-Space Selective Refusal Restoration
Abstract
Prefill jailbreaks can cause safety-aligned language models to continue an unsafe assistant prefix even when they would refuse the same request under normal decoding. We study how the model's internal safety state evolves along this attacker-controlled trajectory. Using the Jacobian lens, we track vocabulary-grounded harmfulness and refusal signals from the assistant header through prefilling and subsequent generation. We identify a characteristic phenomenon, which we call refusal-state erosion: both the refusal signal and the harmfulness signal weaken as the prefill progresses, but refusal progressively fades across the prefill trajectory, while harmfulness-state remains detectable after the prefill. Motivated by this observation, we introduce J-Anchor, a training-free inference-time defense that uses the model's pre-attack safety state as a request-specific anchor and selectively restores weakened refusal in J-space during prefilling and decoding. J-Anchor reduces fixed-prefill attack success from 54–89% to at most 1.0% across different model families. It reduces deep-prefill attack success from 83–94% to at most 2.5%. J-Anchor largely preserves benign behavior and reasoning ability, and ablations show that its effectiveness depends on restoring refusal in the J-space refusal subspace rather than applying generic activation perturbations. Our results suggest that preserving latent refusal dynamics offers a promising route to robust defense against prefill jailbreaks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.