TFL: Targeted Bit-Flip Attack on Large Language Models
Abstract
Large language models (LLMs) are increasingly deployed in safety and security-critical applications, raising concerns about their robustness to model parameter fault injection attacks. Recent studies have shown that bit-flip attacks (BFAs), which exploit computer main memory vulnerabilities through flipping a small number of bits in model weights, can severely disrupt LLM behavior. However, existing BFA on LLMs largely induce un-targeted failure or accuracy degradation, offering limited control over manipulating specific or targeted outputs. In this paper, we present TFL, a novel targeted bit-flip attack framework that enables precise manipulation of LLM outputs for selected prompts while limiting degradation on unrelated inputs. Within our TFL framework, we propose a novel keyword-focused attack loss to promote attacker-specified target tokens in generative outputs, together with an auxiliary utility score that balances attack effectiveness against collateral performance impact on benign data. We evaluate TFL across factual QA, summarization, code generation, and mathematical reasoning, using multiple LLMs, including Qwen, DeepSeek, and Llama. Across the main experiments, TFL induces targeted output changes with only 1–9 weight bit flips. It also incurs substantially less collateral degradation than degradation-oriented BFA baselines. These results demonstrate that small weight perturbations can selectively alter LLM content and behavior, exposing a novel threat beyond untargeted model degradation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.