Hidden in Meaning: Sense-Specific Backdoors in Large Language Models
Abstract
Large language models (LLMs) excel at language understanding and reasoning through their sensitivity to semantics, which enables them to capture subtle meaning differences across diverse contexts. However, this capability may also introduce an overlooked security risk: malicious behaviors can potentially be conditioned on semantic distinctions rather than explicit surface triggers. Prior studies on LLM backdoor vulnerabilities mainly focus on triggered by explicit lexical or structural patterns, implicitly assuming malicious activation is surface identifiable. We challenge this assumption by revealing polysemy as a new and stealthy threat surface, where specific word senses can serve as covert triggers that activate malicious behavior only under the target sense while remaining inert otherwise. To systematically investigate this risk, we propose Sense-Aware Backdoor attack (SAB), a model editing framework that combines sense discrimination with selective editing to isolate a discriminative sense subspace and construct editing channels for target injection and benign preservation, achieving activation selectivity with limited data. Extensive experiments across four public benchmarks and a self-constructed benchmark CW20 show that SAB achieves a high attack success rate under the target sense while maintaining minimal to zero activation on non-target senses. Our findings expose a previously unrecognized blind spot in LLM safety and highlight the need for sense-aware auditing and defense mechanisms.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.