Self-Supervised Continuous Dynamics Learning
Abstract
Temporal modeling is essential for general intelligence to understand change and anticipate how the world evolves. Consecutive frames in a video often share similar semantics while carrying different state information. The dynamic state-transition information between these frames is important for describing their evolution along a continuous time axis. However, conventional self-supervised methods primarily learn associations between frames, lacking explicit awareness of continuous state evolution. To address this issue, we introduce TemDy, a self-supervised continuous dynamics learning framework comprising Temporal Variation Encoding and Continuous Dynamics Modeling. Temporal Variation Encoding uses 3D rotary position encoding to incorporate spatial and temporal information. A DELTA register learns to encode state transitions through supervision from deterministic state reconstruction and uncertain future prediction. Continuous Dynamics Modeling represents normalized visual states through history-conditioned rotations, separating the direction of change from its temporal progression. Integrating continuous angular velocity from an observation-derived anchor then models temporal dynamics. Evaluations of temporal localization, frame retrieval, and state reconstruction demonstrate that TemDy has improved temporal dynamics modeling capabilities compared with diverse self-supervised baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.