acceptodds
Under review as a conference paper at ICLR 2027

Cross-Encoder Fairness Degradation in Text-Attributed Graphs via Adversarial Node Injection

Abstract

Graph Neural Networks (GNNs) have emerged as the dominant paradigm for modeling graph-structured data. Despite their wide adoption, recent studies reveal that GNNs are vulnerable to adversarial perturbations, particularly fairness attacks. These attacks can induce discriminatory outcomes and cause severe societal consequences in human-centered applications. Consequently, existing research has investigated fairness attacks targeting GNNs on traditional graphs to expose potential risks and ultimately guide the design of more robust systems. However, to the best of our knowledge, no existing fairness attack methods target GNNs processing Text-Attributed Graphs (TAGs), even though TAGs are ubiquitous and critically important in semantically rich domains like social networks. As a result, the community currently lacks a fundamental understanding of the fairness vulnerabilities of GNNs when processing TAGs. To bridge this gap, we propose a Universal Fairness Poisoning Attack against GNNs on TAG (UFPAT), which induces fairness-bias in the victim GNN by injecting malicious nodes. UFPAT achieves the first attack of generalizing across different text encoders. It leverages Large Language Models to refine the textual features of malicious nodes, ensuring high readability for real-world feasibility. We benchmark UFPAT against SOTA methods on three real-world TAG datasets. Results show that UFPAT is the only method capable of launching fairness poisoning attacks targeting GNNs on TAGs while generating readable text. It generalizes cross-text-encoders and attacks both standard and fairness-aware GNNs. This work reveals key fairness vulnerabilities in GNNs for TAGs and offers insights for improving robustness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.