EvUBody: Resolving Kinematic Ambiguity in Event-Assisted Joint Upper-Body and Hand Reconstruction
Abstract
Extracting precise human manipulation trajectories from video is critical for scaling embodied robot learning. While sparse RGB frames frequently miss high-frequency kinematics and suffer from motion blur, event streams have emerged as a powerful modality to capture asynchronous dynamic changes. However, existing event-based paradigms are overwhelmingly restricted to isolated hand motion, failing to account for the macro upper-body movements that induce complex, ambiguous event patterns across the hand during real-world manipulation. Transitioning to a joint upper-body and hand framework is strictly necessary, yet it has been severely bottlenecked by a fundamental data collection conflict: obtaining precise, continuous 3D joint annotations typically requires invasive wearable sensors or dense optical markers, which destroy the natural visual appearance of the human and suffer from severe occlusion during hand-object interaction. Furthermore, merging separated body and hand tracking systems introduces persistent spatial discontinuities at the wrist. To solve this, we introduce EvUBody, the first comprehensive dataset and benchmark that successfully balances natural RGB-Event visual observation with high-fidelity, spatially aligned 3D kinematic references for the coupled arm-and-hand system. Comprising both simulated and real-world subsets, EvUBody enables our proposed causal reconstruction framework, which fuses sparse RGB spatial anchors with continuous event streams via joint spatio-temporal modeling. By evaluating across varying frame rates and lighting conditions, we demonstrate that jointly modeling the upper body resolves event ambiguities, allowing our framework to recover true manipulation trajectories that RGB-only priors and hand-isolated approaches fail to capture.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.