acceptodds
Under review as a conference paper at ICLR 2027

MoTween: In-Betweening for Part-wise Text-to-Motion Generation

Abstract

Text-to-motion generation aims to synthesize semantically consistent 3D human motions from natural language descriptions. Although recent approaches have achieved remarkable progress, existing methods either overlook the inherent hierarchical structure of the human body or struggle to capture the temporal dynamics described in the text. To address these limitations, we propose MoTween, a novel two-stage framework consisting of Hierarchical Vector Quantized Variational Autoencoder(Hierarchical VQ-VAE) and In-Betweening Motion Generation. In the first stage, we introduce a Hierarchical VQ-VAE that decomposes each pose into fine-grained body parts and coarse body groups, propagating group-level context into part representations to reflect the inherent body structure. In the second stage, we generate motion representations at semantically critical keyframes and subsequently complete the full sequence by in-betweening the remaining frames, enabling temporally coherent generation. Experiments on the HumanML3D and KIT-ML datasets demonstrate that MoTween outperforms existing state-of-the-art methods on both benchmarks, while also enabling flexible part-level motion editing. Code and models will be released soon.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.