Robust to the Wrong Attacker: Separating Model from Attacker in Graph Transformer Robustness
Abstract
Graph-transformer robustness comparisons can use edge edits chosen by a linearized GCN surrogate. Poor surrogate fit can understate vulnerability and reverse model rankings. We introduce a paired, reference-relative estimand: the surrogate's shortfall against a stated strong attack on the victim minus the corresponding shortfall on a named GCN baseline, evaluated on common targets and paired over ten seeds. Our contribution is a controlled evaluation protocol and paired estimand for measuring surrogate-induced bias; the attack procedure is deliberately conventional. On Cora and Citeseer, linearized-surrogate GCN-to-transformer gaps are 16 to 24 points versus 4 to 13 under full-access greedy search. The excess is positive in all eight victim-graph cells against a plain GCN and separated from zero in six; against an accuracy-matched GCN it is positive in six and separated above zero in five. In three of these sixteen comparisons the surrogate inverts the ranking. On Pubmed, at nearly equal clean accuracy, the excess is 9.9 and 9.4 points against the two GCNs, both separated from zero. An exhaustive per-step candidate reference yields positive estimates separated from zero for the GPS-style transformer against both baselines on both graphs. No excess is detected for the official GPSConv layer on Cora or for an SGFormer-style victim, so the claim concerns surrogate fit, not attention. With an attacker-owned shadow and at most 50 score queries per target, our diagnostic falls no more than 5 points below greedy white-box search across 16 complete-record conditions and sometimes exceeds it; in the official pipeline of Foth et al. (2025) it flips 61 percent of targets against 63 for their adaptive white-box attack. The protocol also separates defences: adversarial training lowers attack success, whereas a Jaccard edge filter mainly reduces transferability. Of its apparent 32.3-point protection against blind transfer on Cora, filter-aware transfer recovers 24.7 without victim weights or victim queries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.