SWIFT: Sequence-aware Wasserstein Informed Flow Transport for Conditional Audio Generation
Abstract
Existing methods for coupling audio sequences with noise sequences in flow matching use a frame-level Euclidean ground metric that treats each sequence as a flat vector, entirely ignoring both the temporal structure of acoustic events and the semantic content of the conditioning signal. We propose SWIFT (Sequence-aware Wasserstein Informed Flow Transport), a flow matching training framework for audio generation built on a novel ground metric called the Sliced Wasserstein Sequence Distance (SWSD). SWSD lifts each audio and noise sequence to a probability measure on its time-augmented path embedding, encoding temporal ordering explicitly into the geometry of the ground metric, and computes sliced Wasserstein distances using a semantic-aware slicing measure that concentrates projections along the acoustic subspace indicated by the text or visual conditioning vector. We validate SWIFT on two audio generation tasks: text-to-audio and video-to-audio. SWIFT outperforms baselines across evaluation metrics and inference budgets, with the largest improvements at low numbers of function evaluations, consistent with the design motivation that a semantically informed coupling produces straighter flow paths consistent with the test-time prior. We further demonstrate that a model pretrained with SWIFT provides a stronger initialisation for the reflow procedure, amplifying the benefit of trajectory straightening and achieving lower Fréchet Audio Distance at a fixed inference budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.