QuanTopo: Topology-Directed Hierarchical Text-to-Motion Generation
Abstract
Text-to-motion generation has advanced in realism and text alignment, yet composite descriptions require global sequence organization alongside diverse, valid realizations. We introduce QuanTopo, a topology-directed hierarchical framework that learns the discrete branch of a hybrid generator as an explicit motion plan rather than only a reconstruction code. QuanTopo addresses a target–use gap: a useful plan needs both an objective that specifies what it should preserve and a structured role in synthesis. To define plan content, persistent homology summarizes multiscale relations along a trajectory of pose, root motion, and local change, emphasizing connectivity, revisitation, and cycles beyond framewise fidelity. To make this content actionable, the plan establishes anatomical context before continuous realization enters joint synthesis, while the continuous pathway models alternative timings, postures, and coordination patterns. A predictor conditioned on text and supplied duration connects the same trajectory target to inference. QuanTopo obtains FID scores of 0.027/0.122 on HumanML3D/KIT-ML and reaches 0.804 composite-action sequence satisfaction on a prespecified subset of composite HumanML3D captions, versus 0.782 for the strongest baseline. In this hybrid setting, ablations and interventions support aligning a plan's information target with its role in synthesis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.