When Do Heterogeneous Text-Attributed Graphs Need Graphs?
Abstract
Heterogeneous text-attributed graphs (HTAGs) combine rich node text with typed graph structure. However, existing HTAG evaluations typically fix a single text encoder, leaving unclear how representation changes affect different structural gains. We introduce MeHTAGs Bench, which pairs six HTAG datasets with 15 node representations, including a random control, yielding 90 settings with fixed graphs, labels, target nodes, and splits. Across representation and architecture sweeps, homogeneous message-passing gains generally contract as text representations strengthen, whereas additional relation-aware gains show no consistent contraction. Moreover, the best text-only representation need not be the best graph-model input. We further introduce Text–Graph Alignment (TGA), a label-free, training-free rank statistic that compares the cosine similarities of meta-path-connected target-node pairs with those of randomly sampled pairs. TGA exhibits the strongest absolute rank correlation with residual homogeneous message-passing gains among the evaluated label-free statistics in each of the six HTAGs. Across all six HTAGs, fixed-weight aggregation probes further associate higher input TGA with smaller aggregation-induced alignment gains, even after ceiling normalization. These results show that graph utility depends on the graph–representation pair, with TGA providing an empirical diagnostic of residual homogeneous gains within a fixed HTAG.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.