acceptodds
Under review as a conference paper at ICLR 2027

When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models

Abstract

Graph Language Models (GLMs) have become a promising direction for adapting Large Language Models (LLMs) to graph learning tasks. By transforming graph topology and node information into graph tokens, GLMs allow LLMs to jointly process structured graph inputs and textual instructions. Yet, it remains unclear how the internal behavior of these tokens relates to the task information they contain and their contribution to prediction. In this work, we analyze how LLMs process graph information through graph-token behavior in two representative GLM architectures, LLaGA and TEA-GLM, across seven LLM backbones and seven datasets. **Findings.** We find that large graph-token activations do not reliably indicate greater contribution to prediction. Graph sink tokens emerge as activation-level outliers, with massive activation values along a small set of hidden-state dimensions, but do not necessarily attract the largest attention weights from query tokens. Their positional patterns differ between architectures: LLaGA sinks are overrepresented among padding tokens and rarely occur at target-node positions, whereas TEA-GLM shows no consistent positional preference across backbones and datasets. Pruning and activation patching generally affect LLaGA's node-classification predictions less when applied to sink tokens than to matched non-sink tokens. This distinction is weaker in link prediction and TEA-GLM, while swapping sink and non-sink input embeddings usually causes small accuracy changes in both architectures. Probing further shows that later LLaGA sink states contain node-label information that a simple classifier can read. **Implications.** Together, these results reveal a decoupling between activation-level saliency and graph-semantic utility. Large activations, readable label information, and predictive contribution do not necessarily coincide in the same graph tokens. This distinction provides a basis for assessing graph-token construction, placement, and alignment mechanisms beyond activation magnitude alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.