Observation-Adaptive Spatiotemporal Positional Encoding: A State-Space Perspective
Abstract
Learning to perceive and predict the evolving physical world is central to physical intelligence, but developing intrinsic spatiotemporal perception remains a key challenge. Transformers represent position through positional encoding (PE). However, temporal coordinates in PE typically use raw time indices, which specify when observations are sampled but do not reflect how much the system state changes between them. To address this limitation, we propose SpatiotemporalPE, which combines spatial structure and temporal evolution. Using a state-space formulation, we relate changes in the underlying state to elapsed time. Specifically, we estimate local state variation from observation differences through averaging and noise correction to construct Observation-difference Coordinate. We instantiate SpatiotemporalPE for molecular dynamics simulation and video understanding. SpatiotemporalPE achieves the best MAE on 8 of 10 molecular dynamics tasks and reduces average MAE by 14.2% relative to the best baselines. SpatiotemporalPE improves overall accuracy from 43.85% to 44.63% in Video experiment. These results support observation-adaptive coordinates as a cross-domain inductive bias for spatiotemporal modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.