PHE-DiT: Phase-aware Hierarchical Energy Conditioning for Flexible-Length Dance-to-Music Generation
Abstract
Dance-to-music generation seeks to synthesize music that is temporally synchronized with human dance movements across multiple scales, requiring global tempo consistency, beat-level correspondence, and onset-level precision. However, existing methods typically condition on coarse motion features without explicitly modeling the hierarchical rhythmic structure of dance, and are constrained to fixed-length generation. We propose PHE-DiT, a conditional flow matching framework with a cross-modal transformer backbone that addresses these limitations through three key designs. First, phase-aware input fusion encodes instantaneous phase extracted from dance motion via the Hilbert transform as sinusoidal features and injects them into the input latent sequence, providing an explicit beat-cycle clock signal complementary to Rotary Position Embedding. Second, hierarchical split energy conditioning decomposes motion energy into slow and fast components via Butterworth filtering and applies them to earlier and later transformer layers via cross-attention, respectively, achieving coarse-to-fine alignment from beat-level to onset-level. Third, variable-length sequence support with length-aware attention masking enables flexible-length music generation from short clips to full-length choreographies within a single unified model. Extensive experiments on AIST++ and TikTok benchmarks demonstrate that PHE-DiT achieves state-of-the-art performance in rhythmic alignment, onset precision, and music quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.