AnyAudio: Accurate, Versatile, and Training-Free Audio Control via Steering Generation
Abstract
Text-to-audio generation systems have been capable of generating high-fidelity samples from natural language prompts with a scalable diffusion transformer (DiT). Recent advances have further allowed fine-grained controllability, which mainly spans precise timing control, long-form audio generation, as well as acoustic control (e.g., pitch and loudness). However, existing methods usually leverage conditional learning, either from scratch or fine-tuning, or training-freely explore latent manipulation, both of which hinder them to achieve versatile audio control efficiently. In this work, we present a holistic audio control framework, AnyAudio, which precisely steers the sampling trajectory to desired target, fully exploiting the advantages of pre-trained audio DiT. Specifically, we first design Timing Attention Control in the condition space, aligning each sound event with their designated time intervals for precise event timing. Furthermore, we present Latent Spectral Decoupling in the diffusion intermediate representation space, fusing overlapping segments in the frequency domain for coherent long-form generation. Then, we develop Reference Attribute Control in waveform space, matching differentiable reference attributes for fine-grained physical control. These cross-space control signals share a unified paradigm of differentiable matching loss optimized via gradient descent, which can be integrated into a diffusion sampling trajectory in an accurate, composable, and training-free manner. Extensive experiments demonstrate that AnyAudio achieves state-of-the-art performance in control of timing, long-form generation, and acoustic characteristics, independently and jointly, without training resource. Demo Page: https://audioguidance.github.io/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.