Don’t Describe, Demonstrate: Learning Transferable Music Edits from Before–After Examples
Abstract
Existing music editing methods mainly rely on text instructions, which are often insufficient for fine-grained or difficult-to-describe transformations. We introduce editing by demonstration, where a before–after music pair specifies an edit to be transferred to unseen inputs. This setting poses two key challenges: (1) translating a demonstrated transformation into an executable editing condition, and (2) isolating the transferable edit from reference-specific musical content. To address them, we propose Edit Condition Synthesis, which recovers a soft edit condition in the native conditioning space of a frozen flow-based editor. We further introduce a transition-aware semantic constraint that captures source-to-target changes and improves the transferability of the synthesized condition. Experiments on instrument and style editing show improved edit transfer and source preservation over existing text-guided and demonstration-based methods. An online demo with audio examples is available here.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.