GINO: Graph DINO through Canonical Token Spaces
Abstract
Modern self-supervised learning in language and vision leverages tokens with well-defined positional structure. Words follow an ordered sequence, while image patches occupy a spatial grid. Graphs, in contrast, have no intrinsic node order, causing any positional machinery tied to input indices to fail under different node labelings. We introduce GINO, a canonicalization-based framework that equips graph tokens with a consistent spatial layout. GINO extracts a canonical node ordering from the graph's spectral structure and uses it to assign deterministic, two-dimensional coordinates to graph edges. Embedded in this consistent coordinate system, positional encodings and DINO/iBOT-style objectives can operate seamlessly, independent of the initial node labeling. We provide theoretical backing for this framework by guaranteeing the ordering's invariance and bounding its proximity to the exact optimum under a normalized Dirichlet energy. Across molecular benchmarks, self-supervised GINO pretraining improves downstream performance and achieves competitive transfer using only 2D graph structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.