acceptodds
Under review as a conference paper at ICLR 2027

Beyond Node Tokens: Motifs as Structural Context for Graph Tokenization

Abstract

Graph tokenization enables graph-level learning with discrete symbols drawn from a shared vocabulary. A common approach represents each graph as an unordered multiset of quantized node tokens. We analyze the capacity of this representation and empirically unveil the inter-graph token collapse issue, i.e., structurally different graphs can receive highly similar token representations. To remedy this problem, we propose , a graph tokenization framework that uses recurring motifs as reusable structural context. constructs a shared motif vocabulary and learns motif embeddings through soft node matching supervised by motif–graph occurrence, with collision regularization to discourage many-to-one assignments. These embeddings guide node representation learning before quantization, encouraging node tokens to retain both attribute and motif-level information. A Transformer then jointly processes node and motif tokens through token-type-aware attention and motif-conditioned reconstruction of masked node-token embeddings, followed by fine-tuning for graph-level tasks. Empirically, achieves the best mean performance among evaluated methods on almost all classification and regression benchmarks, including a 17.1% reduction in mean absolute error over the strongest evaluated baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.