Human-Aware Video Generation Acceleration for Reconstruction-Consistent Video-to-Motion Synthesis
Abstract
Text-to-video-to-motion synthesis can scale semantic coverage with abundant, diverse 2D videos without requiring paired text–3D motion-capture data at the same scale, but generating the video intermediate is slow. We introduce a training-free, human-aware accelerator for this bottleneck. It identifies a space–time human/action tube from signals already available in a Video DiT, processes most denoising calls only on the video-latent tokens inside the tube, reuses predictions outside it, and periodically refreshes the full frame. The resulting full-frame video and tube are passed to an unchanged motion reconstructor. The routing operator is shared across architectures, while tube estimation follows each model's native signals. On Wan2.2, the selected setting accelerates denoising by 2.4x. Across two architecturally distinct Video DiTs, paired reconstruction and motion-quality tests indicate that the accelerated videos generally retain the local evidence needed for motion recovery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.