Two-Stage Data Poisoning Attack against the Large Language Model for Power Systems
Abstract
Large language models are increasingly being explored as decision-support components in power-system operation, but their dependence on data-driven learning introduces security risks during training and fine-tuning. This paper studies a stealthy data poisoning threat against a Transformer-based language-model-assisted power-system decision pipeline. We propose a two-stage state-aware attack framework. First, a local surrogate model and Jacobian-based augmentation are used to characterize sensitive input dimensions and local decision behavior. A multidimensional vulnerability index is then used to identify critical operating periods, while surrogate self-attention scores are used to prioritize critical nodes. Second, a low-rate poisoning strategy is formulated as a constrained bilevel optimization problem that jointly considers physical feasibility, statistical consistency, and attack effectiveness. An augmented-Lagrangian procedure is used to solve the resulting constrained problem, followed by sparse injections aligned with the system’s operating cycle. Experiments on an IEEE 39-bus simulation with a DeepSeek-R1- Distill-Qwen-1.5B-based decision pipeline evaluate target localization, detection evasion, and physical degradation. The reported results show that state-aware localization concentrates the observed degradation on the selected spatio-temporal targets, while the proposed constrained poisoning remains below the evaluated bad-data and distributional detection thresholds.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.