Looping in Context: Node Classification with Looped Prior-Fitted Networks
Abstract
Prior-fitted networks (PFNs) for graphs classify nodes on arbitrary graphs in a single forward pass, yet how their predictions form across depth is unknown. For tabular foundation models, mechanistic work shows that depth is iterative *refinement*, so one transformer block, looped, recovers a full-depth model. We ask whether this holds for graphs, where every layer also performs one hop of *propagation*. Six mechanistic experiments on the 12-layer NodePFN over 23 benchmarks show that depth plays two roles. Attention transfers the labels in the first layer and is redundant and self-repairing in later layers. Message passing stays necessary on homophilous graphs, where the loss from a removed hop persists to the last layer. The depth at which predictions form therefore grows with homophily, in NodePFN and in a second PFN for graphs, GraphPFN, and a two-class analysis explains why no fixed depth can be optimal across homophily levels. We propose **Looped NodePFN**, a single block applied repeatedly, whose depth is selected per graph *in context* from held-out labeled nodes, with no gradient steps and no extra parameters. With a tenth of the parameters it reaches 66.3% mean accuracy against 66.6% for the 12-layer model, and a two-layer block reaches 66.5%. Once its depth is chosen, the one-layer block predicts up to 5.1 times faster. A PFN for graphs thus needs only one block, but how many times it is applied, and in how many of those iterations it passes messages, should follow the graph.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.