HEDGE: Hierarchical Expressive Dance Generation with Long-term Music Structure Guidance
Abstract
Generating dance from music not only requires precise beat alignment but also synchronization with the global musical structure, especially in long sequences containing multiple sections such as intros, verses, and choruses. Existing methods primarily focus on beat-level alignment, but often fail to capture and reflect broader musical dynamics, which limits expressive variation and results in choreography that does not adapt well to section-level transitions. To address these limitations, we introduce hierarchical segmentation for both music and motion, organizing them into section and interval levels to more effectively capture structural and rhythmic details. Building on this representation, we propose HEDGE, a two-level diffusion-based framework that generates dance sequences conditioned on both local rhythmic cues and global musical structure. Our method produces choreography that is not only realistic and musically synchronized, but also dynamically adapts to the expressive changes across musical sections. Extensive experiments demonstrate that HEDGE achieves superior responsiveness to musical structure, offering a significant advancement toward automated choreography that faithfully reflects the full narrative and expressive intent of music.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.