acceptodds
Under review as a conference paper at ICLR 2027

PHE-DiT: Phase-aware Hierarchical Energy Conditioning for Flexible-Length Dance-to-Music Generation

Abstract

Dance-to-music generation seeks to synthesize music that is temporally synchronized with human dance movements across multiple scales, requiring global tempo consistency, beat-level correspondence, and onset-level precision. However, existing methods typically condition on coarse motion features without explicitly modeling the hierarchical rhythmic structure of dance, and are constrained to fixed-length generation. We propose PHE-DiT, a conditional flow matching framework with a cross-modal transformer backbone that addresses these limitations through three key designs. First, phase-aware input fusion encodes instantaneous phase extracted from dance motion via the Hilbert transform as sinusoidal features and injects them into the input latent sequence, providing an explicit beat-cycle clock signal complementary to Rotary Position Embedding. Second, hierarchical split energy conditioning decomposes motion energy into slow and fast components via Butterworth filtering and applies them to earlier and later transformer layers via cross-attention, respectively, achieving coarse-to-fine alignment from beat-level to onset-level. Third, variable-length sequence support with length-aware attention masking enables flexible-length music generation from short clips to full-length choreographies within a single unified model. Extensive experiments on AIST++ and TikTok benchmarks demonstrate that PHE-DiT achieves state-of-the-art performance in rhythmic alignment, onset precision, and music quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.