Quantifying Graph Structural Memorization in Large Language Models
Abstract
Memorization in Large Language Models (LLMs) is well documented for sequential content such as text and code. As LLMs are increasingly trained with graph-structured data, however, whether memorization also extends to graph structure remains largely unexplored. To study this question, we introduce an exposure-controlled evaluation framework that examines how graph recoverability changes with training exposure. The framework separates training-specific recovery from general graph prediction by comparing exposed target graphs against graphs never exposed during training. Using this framework, we observe an exposure-dependent form of graph structural memorization: graph structure becomes increasingly recoverable as the model is exposed more frequently to the target graph during training. Motivated by this observation, we formalize the underlying recovery procedure as training-graph stealing, a black-box attack against graph-trained LLMs that makes this form of memorization concretely exploitable. Given candidate nodes and their textual descriptions, the attack queries each node pair to infer adjacency and assembles the predicted edges into a reconstructed graph. Experiments show that exposure-dependent recovery persists across graph-access settings, graph scales and model families. Together, these findings demonstrate that LLM memorization can extend beyond sequential content to graph structure, creating a new attack surface for extracting graph-structured training data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.