Forget Less, Generalize More: Unifying Temporal and Structural Adaptation for Dynamic Graphs
Abstract
Representation learning on dynamic graphs requires capturing complex dependencies that evolve across both time and structure. Existing approaches typically adopt fixed temporal decay schemes or predetermined structural propagation depths, limiting their ability to generalize across graphs with diverse interaction frequencies and topological characteristics. We propose Dual-Scale Retentive Dynamics (DSRD), a unified framework that maintains a retentive representation state encoding both temporal memory and structural context. DSRD introduces two key components: (i) a retentive state with dual-scale adaptation that jointly models temporal dynamics and structural propagation within a single recurrent formulation, and (ii) adaptive decay kernels with learnable time-sensitivity parameters that automatically balance short-term responsiveness and long-term retention based on the underlying temporal and structural interaction patterns. We provide theoretical analysis that expands the recurrent state updates into weighted aggregation over temporal walks, together with boundedness guarantees for the learned dynamics. Extensive experiments on 14 real-world benchmarks demonstrate that DSRD achieves state-of-the-art performance on both link prediction and node classification tasks, with strong generalization across transductive and inductive settings. The implementation code is available at [[Code]](https://anonymous.4open.science/r/DSRD).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.