Learning World Dynamics from Deblurring Discrepancy for Hand Trajectory Prediction
Abstract
Predicting human hand trajectories in egocentric views is crucial for mixed reality, visual imitation learning, and human-robot collaboration. It enables autonomous systems to anticipate human intent and support timely interaction. However, this task remains challenging in dynamic interaction scenes, where moving objects and flexible hand adjustments introduce complex interaction dynamics. World models offer a promising way to address this challenge by modeling and predicting future world states beyond hand positions. Nonetheless, pronounced motion blur from fast hand-object motion and rapid viewpoint changes disturbs the observed interaction patterns and makes future visual states highly uncertain. In this paper, we propose to learn world dynamics from ***deblurring discrepancy***, a feature-level residual between blurred and deblurred visual representations. Our key insight is that motion blur manifests as visual ambiguity in image space, but originates from structured temporal integration during image formation. The feature discrepancy induced by deblurring thus preserves motion-related cues of hands, objects, and the camera. Based on this, we propose a deblurring-discrepancy-guided framework, dubbed **DynamicHand**, for egocentric hand trajectory prediction (EHTP). **DynamicHand** extracts blur-induced motion cues from deblurring discrepancy features and predicts their future evolution to facilitate world-state modeling and hand forecasting. We further design a discrepancy-aware distillation strategy, where discrepancy strength and variation adaptively weight the transfer of teacher latents to a lightweight student model for faster inference. Extensive experiments on public datasets and our new benchmark show that **DynamicHand** consistently improves EHTP performance over existing baselines, demonstrating that deblurring discrepancy can serve as an effective bridge between visual restoration and world-dynamics learning in dynamic interaction scenes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.