acceptodds
Under review as a conference paper at ICLR 2027

Graph2Rule: Learning Behavioral Rules from Transformer Computation Graphs

Abstract

Large language models (LLMs) exhibit distinct internal computation patterns across different input types, yet these patterns remain difficult to characterize due to complex interactions among distributed attention heads and feed-forward neurons. Our goal is to understand whether these component interactions contain behavior-specific patterns and whether these patterns can be expressed in an interpretable form. To achieve this, we propose GRAPH2RULE, a framework to represent the computation of a pretrained transformer as a directed weighted graph. In this graph, nodes represent selected attention heads and feed-forward neurons, while edges represent their forward interactions. We begin by testing whether these computation graphs contain sufficient information to distinguish behavioral classes without access to the original input text. Across three datasets and multiple transformer backbones, our graph-based classifiers achieve strong performance. Importantly, the graph representation and GNN classifier enable us to analyze the structure underlying these predictions via GNN explainability techniques. Specifically, we extract sample-level explanation subgraphs, mine recurring class-discriminative motifs from these subgraphs, and convert these motifs into interpretable logical rules. Our results show that transformer component interactions contain structured behavioral information that can be represented as graphs and translated into compact and human-interpretable symbolic rules.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.