Radioactive Watermarking of Graph Node Classification Datasets for Scalable Owner Attribution
Abstract
Graph datasets used in industrial graph neural network (GNN) applications are valuable proprietary assets, often licensed to multiple buyers under restricted-use agreements. When a suspect GNN trained on a buyer's leaked dataset copy appears, the owner must identify the responsible buyer, since only buyer-level attribution allows terminating the license and pursuing legal remedies. However, this task is unsupported by existing dataset watermarks which provide only binary ownership verification and fail under multi-buyer owner tracing. In this paper, we introduce the first radioactive watermarking method for graph node classification datasets that supports scalable owner attribution. Radioactive means that the watermark signal persists through training and is recoverable from model outputs, without access to the released data, the buyer's training procedure, or the GNN architecture. Detection is therefore architecture-agnostic and black-box, meaning it can be applied to suspect models with different architectures, and the underlying test is a closed-form statistic with a provably bounded false-positive rate, jointly protecting dataset and model intellectual property. We attribute a leaked suspect to its true buyer at marketplace scale via a per-buyer watermarking test, with a closed-form attribution rate that saturates with the number of buyers. Across five benchmarks and eight GNN architectures, our method achieves a stealthy yet highly detectable watermark signal that remains robust under cross-architecture detection and four removal attacks, while incurring negligible degradation in task accuracy, and provides reliable owner traceability at the scale of multi-buyer data marketplaces.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.