acceptodds
Under review as a conference paper at ICLR 2027

Which Move, Which Moment? CueDance: Efficient Retrieval-Augmented Music-to-Dance Generation via Choreographic World Modeling

Abstract

People naturally move to music, while choreographic experience helps dancers determine which move fits which moment. This observation motivates us to view music-to-dance generation as two complementary capabilities: continuous responses to music and the selective use of motion experience at key moments. Existing retrieval-augmented methods introduce real motion references, but the temporal allocation of sparse references and sequential refinement dependent on an initial draft still limit their use of memory. We present CueDance, which connects what motion to retrieve with when to introduce it through a choreographic world model. Anticipate–Attribute–Anchor (3A) proposes motion hypotheses from musical transitions, selects reference moments based on explanatory gain (reduction in prediction error), and retrieves real motions as anchors. Music-driven retrieval runs in parallel with flow sampling, while Mid-Flow Anchor Inpainting integrates the references into the same generation pass. A part-aware generator and local-to-global flow supervision further support natural transitions around the references and efficient synthesis. CueDance achieves state-of-the-art dance generation performance on FineDance and AIST++, with a user study and systematic ablations further validating the effectiveness of selective choreographic memory.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.