Robust Spatio-Temporal Adaptation towards 4D Point Cloud Understanding
Abstract
4D point cloud understanding jointly models fine-grained 3D geometry and temporal dynamics, playing a crucial role in robotics, autonomous driving and AR/VR. However, point cloud videos are often degraded by hardware noise, object occlusion and stream disruptions in real-world deployment. We reveal that existing state-of-the-art 4D point cloud understanding methods are vulnerable to such disturbances and formulate a unified evaluation protocol for spatial and temporal corruptions. To address these intrinsic weaknesses, we further propose a novel training-free method, Spatio-TEmporal Adaptation with Decoding dYnamics (STEADY), which adapts corrupted geometric evidence and recovers disrupted temporal dynamics at test time. Specifically, our proposed STEADY method comprises two stages: (1) The Hierarchical Spatial Cache retrieves complementary global and local geometric cues while Soft Temporal Consistency Gate modulates cache corrections based on inter-frame consistency. (2) The Dynamics-Aware Decoder incorporates dynamics of transition and duration priors to recover temporally coherent predictions over the complete sequence. We conduct extensive experiments on three widely used 4D point cloud understanding benchmarks, including MSR-Action3D, NTU-RGBD60 and HOI4D, under eight corruption conditions across both spatial and temporal dimensions. Experimental results demonstrate that STEADY consistently improves the overall robustness of various representative backbones across spatial and temporal corruptions. Our anonymous code is available at https://anonymous.4open.science/r/STEADY/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.