acceptodds
Under review as a conference paper at ICLR 2027

Quantization-Conditioned Backdoors on Large Language Models via Attention Sinks

Abstract

Post-training quantization is widely used to reduce the deployment cost of large language models (LLMs), but recent studies have shown that quantization may be exploited as a backdoor trigger: a released model may appear benign before quantization while exhibiting attacker-specified behaviors after quantization. Existing quantization-conditioned backdoor attacks first implant the target behavior into the quantized model and then restore benign behavior at full precision while keeping the weights within a subspace that yields the same quantized values as the backdoored counterpart. This requires optimizing a complex objective within a highly constrained weight space, leading to limited attack effectiveness. To address this issue, we introduce SinkQCB, which separates the insertion of malicious behavior from the mechanism that triggers it. More specifically, the malicious behaviors are encoded in a small set of attention heads, and attention sinks are utilized to keep these heads suppressed at full precision and activated after quantization. Extensive experiments on commonly considered LLMs, malicious tasks, and quantization methods demonstrate that SinkQCB achieves a higher attack success rate (ASR) after quantization while keeping ASR at or below 1% and maintaining competitive performance on benign tasks at full precision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.