acceptodds
Under review as a conference paper at ICLR 2027

GraphToken: A General Graph Vocabulary for Text-attributed Graph Foundation Models

Abstract

Graph foundation models (GFMs) aim to enable generalizable graph learning across diverse tasks and domains, yet constructing transferable representation units for text-attributed graphs (TAGs) remains challenging due to the coupling of graph structure and textual semantics. Existing graph tokenization methods often derive tokens primarily from structural patterns or rely on continuous graph embeddings, leaving the learning of a shared discrete vocabulary over structure–text patterns underexplored. To address this challenge, we propose **GraphToken**, a motif-based discrete vocabulary learning framework for TAGs. For each target node, GraphToken samples multiple small rooted motif instances from its local neighborhood and jointly encodes their topology and node texts into root-conditioned motif representations. These representations are vector-quantized against a shared learnable codebook, where each codebook entry corresponds to a reusable **graph-token** shared by structurally and semantically related instances. A target node is consequently represented by a common set of graph-tokens that characterize its local structure–text context. Importantly, GraphToken decouples vocabulary induction from downstream prediction, allowing the learned tokens to be used by Graph Neural Networks (GNNs), Transformers, and Large Language Model (LLM)-based approaches alike, regardless of their architectural differences. Extensive experiments across graph domains, tasks, and model architectures demonstrate the effectiveness and generalizability of the learned GraphToken vocabulary.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.