acceptodds
Under review as a conference paper at ICLR 2027

EchoGhost: A Frequency-Structural Backdoor Attack and Spectral-Counterfactual Defense for Graph Neural Networks

Abstract

Graph neural networks (GNNs) are increasingly deployed in security-sensitive domains, yet their message-passing mechanism can amplify hidden training-time perturbations and convert small poisoned patterns into systematic inference-time failures. Existing graph backdoor attacks typically rely on subgraph triggers, node injection, or local feature manipulation, which can leave detectable topological or feature-space artifacts. This article proposes EchoGhost, a frequency-structural backdoor framework for node classification in GNNs. EchoGhost embeds a universal echo trigger into selected node features through controlled frequency-band modulation and reinforces the trigger by injecting similarity-guided ghost edges among low-activity nodes. The resulting poisoned graph causes a trained GNN to map trigger-bearing nodes to an attacker-chosen target class while preserving clean-task utility. Unlike purely structural triggers, EchoGhost couples spectral feature evidence with graph propagation, creating a dual-domain attack surface that neither feature anomaly detection nor topology-only inspection fully captures. To study and mitigate this threat, we introduce a Spectral-Counterfactual Trigger Verification defense that identifies suspicious training nodes using robust spectral-energy statistics and verifies causal trigger dependence through counterfactual de-echoing. We sanitize confirmed nodes by removing them from the training set, then retrain from scratch on the original clean graph. We evaluate EchoGhost and the proposed defense on four benchmark graph datasets, such as CoauthorPhysics, CoauthorCS, Citeseer, and Cora, using three GNN architectures: GCN, GAT, and GraphSAGE, reporting all results as mean standard deviation over five random seeds. Across all 12 configurations, EchoGhost achieves an average attack success rate (ASR) of 98.04% with an average clean accuracy drop of 7.22%. The defense mechanism reduces the average ASR to 14.52% (an 85.21% relative reduction) while maintaining a false positive rate of 0.01% across all settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.