STAGE: Scaling Novel Graph Generation via Lightweight Structure-Guided Autoregressive Models
Abstract
Generating realistic and diverse graphs is a fundamental challenge because molecules, materials, circuits, and malware can all be represented as graphs. A useful generator must produce novel graphs that follow the underlying data distribution rather than simply reproduce training examples. For instance, regenerating training molecules yields no new drug or material candidates, while replicating known malware graphs provides no new variants for improving detectors. Graph generators must therefore scale to large, sparse graphs while producing novel, diverse, and distributionally consistent samples. Current methods fall short on at least one requirement. Dense autoregressive and diffusion models incur at least quadratic cost in graph size, and sparse autoregressive models scale but, like dense ones, fit the training distribution, so novelty arises only incidentally. To our knowledge, no existing method achieves both scalability and novelty. We introduce STAGE, a lightweight autoregressive graph generator designed for both scalability and novelty. For scalability, STAGE serializes graphs as flat bit-level streams of lexicographically sorted edges, yielding generation with any causal sequence backbone; to make this compact stream learnable, STAGE orders nodes by a breadth-first traversal that enqueues neighbors by SIR-GN structural rank, resolving many of the arbitrary choices left by plain BFS. For novelty, STAGE uses a two-phase training strategy that combines exploration-oriented edge perturbations with self-training on generated samples. Samples are retained either by a label-free Gaussian-mixture density in graph-embedding space or, when available, by a domain validity oracle, steering exploration toward plausible regions. With label-free filtering, STAGE achieves the highest joint validity–uniqueness–novelty (VUN) on QM9, QM7x, Transition1x, and ANI-1x (56.8% vs. 55.0% on QM9; 86.4% vs. 59.4% on QM7x), matches the strongest baseline on MOSES, and remains competitive on ZINC250K. With oracle filtering, it reaches 60.9% VUN on QM9 and 95.1% on MOSES, compared with 55.0% and 85.4% for the strongest baselines. Ablations show that ordering is critical, with random ordering sharply degrading validity. Finally, STAGE trains on MalNet-Tiny function-call graphs with a median of 2,272 nodes, a regime in which DiGress, DeFoG, and SparseDiff exhaust GPU memory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.