Efficient 3D Motion Generation from Visual Demonstrations
Abstract
Existing 3D motion generators face a trade-off between convenient but coarse text-based control and precise but manually specified pose guidance. Moreover, diffusion-based methods further incur substantial iterative sampling costs. We propose FlowMimic, an efficient framework for generating complete 3D human motion from sparse, timestamped 2D pose observations extracted from videos or image sequences. FlowMimic consists of a Motion Latent Codec (MLC), a Time-Aware Pose Encoder (TAPE) for sparse spatiotemporal conditioning, and a Conditional Motion Flow (CMF) that performs Flow Matching directly in the learned motion latent space. The resulting flow supports single-step motion generation while retaining effective control from sparse pose observations. Experiments on AIST++ show that FlowMimic faithfully follows sparse pose observations while maintaining realistic motion quality, and substantially improves inference efficiency over iterative diffusion-based generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.