PrisonBreak: Jailbreaking LLMs with at Most Twenty-Five Targeted Bit-flips
Abstract
We present a new vulnerability in aligned large language models (LLMs): their alignment can be bypassed by flipping only a few bits in their parameter values. We exploit this vulnerability with an attack that jailbreaks billion-parameter LLMs with just 5–25 bit-flips—up to 40× fewer than prior bit-flip attacks. To this end, we present an efficient bit-selection algorithm that identifies critical bits for jailbreaks up to 20× faster than exhaustive strategies. Unlike prompt-based jailbreaks, our method directly alters model behavior in memory at inference time, enabling harmful outputs without any input-level modifications. We evaluate our attack on 10 LLMs, achieving attack success rates (ASRs) of 80–98% with negligible impact on model utility. We further demonstrate an end-to-end exploit via Rowhammer-based fault injection, reliably jailbreaking five models (69–91% ASR) on a GDDR6 GPU by exploiting just two vulnerable physical bit locations. Moreover, bit-flip locations identified in one fine-tuned model transfer to others that share the same pre-trained LLMs, demonstrating the black-box exploitability in standard pretraining-then-finetuning paradigms. We evaluate potential countermeasures and find that our attack remains effective against defenses at various stages of the inference pipeline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.