What Neighbors should an LLM See? A Systematic Empirical Study of Graph Neighborhood Serialization
Abstract
Large language models (LLMs) have recently been used to reason over graph-structured data for graph tasks, including Node Classification (NC) and Link Prediction (LP). Since graph neighborhoods are non-sequential and can span multiple hops beyond the LLM's sequence budget, only a subset of the graph context can be serialized. This motivates us to investigate: What neighbors should an LLM see? We present a systematic empirical study of graph neighborhood serialization on both NC and LP tasks across four datasets: Cora, Citeseer, PubMed, and ogbn-arxiv. We compare random selection, structural ranking, semantic similarity, and task-aware signals and show that no universal neighborhood serialization method is dominant. The effectiveness of structural signals varies between datasets and tasks, with random selection performing competitively in some settings. To analyze the effect of class-relevant context, a label-informed Oracle, used only for diagnostics, improved NC performance in all datasets, with the largest improvement on ogbn-arxiv, where accuracy increased from around 76% to 92%. Task-aware signals, however, revealed different task-related limitations: NC class-relevance signals are only approximate, while LP pair-aware signals are exact when available but often sparse. Cross-serialization analysis showed that the models can adapt to their training-serialization distribution, therefore affecting their performance on different test-time serialization sequences. Overall, our study revealed that the effectiveness of neighborhood serialization depends on several factors and features that are both dataset and task-dependent.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.