acceptodds
Under review as a conference paper at ICLR 2027

THE PRICE OF FORGED EVIDENCE IN SIMILARITY-BASED DEFENSES FOR GRAPH NEURAL NETWORKS

Abstract

Similarity-based defenses for graph neural networks decide whether an edge is genuine by comparing the features of its endpoints. If the attacker can also write those features, that evidence can be forged. Adaptive attacks are known to break such defenses; what has not been measured is how cheap the forgery is and how much of the damage it explains. We price it exactly: for Jaccard and cosine filters we give the fewest feature changes that make an inserted edge pass, in closed form, and about one copied word for each inserted edge the filter would remove makes GCN-Jaccard keep every edge that Metattack or PRBCD inserts, with no change to the attack, for at most 0.46% of the non-zero feature entries on the citation graphs (1.31 to 5.84% on the heterophilous ones). The defense then does harm: on Cora, 136 copied words added to Metattack’s graph leave GCN-Jaccard losing more accuracy than an undefended GCN (15.7 against 13.3 points). We also measure forgery edge by edge: we count which inserted edges each defense keeps, and rerun the defense on clean features while its model still trains on poisoned ones, which separates blinding from ordinary feature damage that no edge filter was built to stop; the pruning counts this literature reports conflate the two. Every countermeasure we test that reads raw features, a higher threshold, a rule that discounts shared words, an edge-aware corroboration check, and GNNGuard’s stricter cosine, only raises the price, while the defender pays in clean edges. Defenses that learn their evidence first are the exception: on Cora and CiteSeer the same copies change neither a STABLE-style purifier nor SimP-GCN by more than 2.8 points, though a large feature budget still blinds them. A joint bilevel attack, CLASP, finds the forgery without seeing the defense. The pattern holds across six graphs spanning homophily 0.22 to 0.81, wherever the attacker’s edges do damage.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.