acceptodds
Under review as a conference paper at ICLR 2027

Learning Pose-Pixel Synergy: Joint Egocentric Video and 3D Hand Pose Generation from a Single Image

Abstract

Predicting future actions based on the current state from an egocentric (Ego) perspective is vital for applications such as embodied AI. Previous works typically predict sparse hand poses or trajectories, or generate action videos relying on restrictive conditioning inputs. We observe that the two forms of Ego actions (i.e., hand pose and video) exhibit inherent synergy: poses provide abstract high-level structured dynamics, while videos preserve detailed pixel-level visual contexts. Inspired by this, we introduce an Egocentric Pose-Pixel Flow Network (EgoPFN), a unified flow-matching-based framework to jointly generate future Ego 3D hand poses and videos from a single image. To overcome the challenges of insufficient visual cues and the pose-pixel representation gap, we first develop a Geometry-aware Spatial-Temporal Propagation Module (GSTPM). It leverages geometric cues from the estimated first-frame depth map, which are integrated into visual features via spatial enhancement and temporal modulation, enabling explicit scene layout perception and persistent geometric guidance. Then, we propose a Pose-Pixel Synergy Module (PPSM), which learns pose flow from observation-anchored intermediate features and injects the learned structured pose information back into video latents via adaptive layer normalization, thereby enhancing the consistency between the generated poses and videos. Extensive experiments on our newly curated benchmark demonstrate that our EgoPFN outperforms many state-of-the-art video generation and hand pose prediction approaches by a large margin. Code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.