Learning Substructure Regularities from Whole-Graph Sequences
Abstract
Autoregressive Transformers trained on whole graphs can reproduce substructure-level regularities without explicit substructure supervision. We investigate what this agreement reveals about substructure generalization, using PCQM4Mv2 as the main testbed. On PCQM4Mv2, under pure ancestral sampling, a Transformer trained on DFS-code serializations reproduces the support ranking of frequent training substructures, even after exact training-string matches are excluded. Finite-context -gram models also achieve substantial aggregate agreement, and on five small TUDataset benchmarks agreement coincides with replay, so this observation alone does not identify how substructures are generated. We therefore remove every training graph containing a specified target and retrain Transformers on an alternative serialization of node and edge additions. Under unconditional sampling, models generate structurally valid graphs containing withheld cyclic and acyclic targets across all three training seeds. The rare outputs containing a withheld carbon triangle are not labeled-isomorphic to any graph in the original, unfiltered training set and cannot be produced by generators restricted to recombining retained training fragments only through bridge connections. These results establish both support reproduction and the generation of specific structures absent from training, without identifying an internal mechanism or a general construction rule. Combining support evaluation, sequence baselines, and controlled target removal provides an empirical framework for characterizing observable substructure generalization and for future investigation of its underlying mechanisms.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.