A Structural Audit of Missing Propagated Evidence in Temporal Link Prediction
Abstract
Local message-passing graph neural networks are the default tool for temporal link prediction, the task of forecasting which node a source will interact with next. Such a model can use source information in scoring a candidate only if that information reaches the candidate's representation, which raises a question rarely asked before training: how much of a benchmark's evaluation set carries source-to-target propagated evidence at all? Expressiveness results describe what a model class can represent once the right input is present, not whether a given query supplies it, and post-hoc structural analyses correlate graph statistics with accuracy over whole datasets, after the cost of training has been paid. We introduce a training-free audit, computed from the history graph by breadth-first reachability and Weisfeiler–Leman refinement, that flags queries whose target lies outside the source's hop ball and, for anonymous models, queries whose target is indistinguishable from a reachable distractor. We prove what the flag says and what it does not: the target representation of a flagged query carries nothing propagated from the source, yet the flag cannot bound the accuracy of the class, because free node embeddings achieve a perfect ranking on any pairwise-consistent query set however many queries are flagged. The audit therefore measures available evidence rather than capacity, and what the missing evidence costs a trained model is an empirical question. On featured graphs, every architecture we train, from local aggregation and global mixing to source-conditioned path reasoning and temporal memory, collapses toward random on flagged queries while performing well on the complement, and the collapse survives matching on degree, depth and width. On feature-free graphs the same gap is largely degree composition; we state the conditions under which the audit is informative, test them forward on held-out graphs, and isolate the one condition under which a non-message-passing baseline recovers the flagged queries and a change of model class is warranted. A graded hop score turns the flag into a pre-training rule for where source conditioning pays off and where it does not.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.