acceptodds
Under review as a conference paper at ICLR 2027

Human-Aware Video Generation Acceleration for Reconstruction-Consistent Video-to-Motion Synthesis

Abstract

Text-to-video-to-motion synthesis can scale semantic coverage with abundant, diverse 2D videos without requiring paired text–3D motion-capture data at the same scale, but generating the video intermediate is slow. We introduce a training-free, human-aware accelerator for this bottleneck. It identifies a space–time human/action tube from signals already available in a Video DiT, processes most denoising calls only on the video-latent tokens inside the tube, reuses predictions outside it, and periodically refreshes the full frame. The resulting full-frame video and tube are passed to an unchanged motion reconstructor. The routing operator is shared across architectures, while tube estimation follows each model's native signals. On Wan2.2, the selected setting accelerates denoising by 2.4x. Across two architecturally distinct Video DiTs, paired reconstruction and motion-quality tests indicate that the accelerated videos generally retain the local evidence needed for motion recovery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.