SceShortcut: Efficient Scene-Aware Motion Synthesis via Shortcut Modeling
Abstract
Previous works on scene-aware human motion synthesis have mainly focused on motion-scene compatibility, with less attention to runtime efficiency. In this work, we present SceShortcut, a shortcut-based framework that accelerates scene-aware motion synthesis through one-step or few-step sampling while producing plausible and scene-compatible motion. To preserve fine-grained scene adaptation under a limited sampling budget, we introduce a global-to-local architecture with two stages trained jointly end-to-end. Specifically, at each sampling step, conditioned on the text instruction and global geometric-semantic scene features, the global stage produces intermediate motion features and predicts a root transport velocity. This velocity yields a coarse clean root motion estimate for querying fine-grained local scene geometry from a dense height field. The local stage incorporates text conditioning and the queried geometry into these intermediate features to predict transport velocities for the remaining motion components and a residual correction to the global-stage root transport velocity. To improve coarse root motion estimates and thereby enhance generation quality, we augment shortcut training with additional root-transport-velocity supervision and root-residual regularization. On TRUMANS and HUMANISE, one-step sampling with SceShortcut achieves approximately speedups over recent advanced open-source baselines while producing plausible motion with comparable motion-scene alignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.