IUC: Intersection–Union Correction for Quantization-Activated Attacks on LLMs
Abstract
While quantization is a key technique for the efficient deployment of large language models (LLMs), recent studies have shown that it may also introduce new security risks: a model that behaves normally in full precision may exhibit malicious behaviors after quantization, such as advertisement injection, over-refusal, or jailbreak, posing a practical threat to the secure deployment of LLMs. To address this issue, we propose **I**ntersection–**U**nion **C**orrection (IUC), a *one-shot* defense method applied directly to the full-precision model. IUC models local parameter intervals under different quantization configurations, where the intersection of these intervals characterizes the shared local reference region underlying the correction, while the union constrains the magnitude of parameter correction. *A smaller intersection indicates a narrower shared local reference region,* and IUC prioritizes the correction of such parameters. With only a single correction applied to the full-precision model, IUC supports secure deployment under multiple commonly used quantization configurations. We conduct experiments across multiple LLMs, different attack scenarios, and representative quantization configurations. The results show that IUC effectively suppresses the activation of malicious behaviors after quantization with low resource overhead, while largely preserving model utility, providing a one-shot pre-sanitization solution for secure model deployment under commonly used quantization configurations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.