acceptodds
Under review as a conference paper at ICLR 2027

One-Step Is Enough to Judge: Hybrid Noise Skimming for Full-band Text-to-Audio Generation

Abstract

Text-to-audio (TTA) systems based on latent diffusion and flow matching achieve high fidelity in the full-band scenarios, but their iterative sampling nature hinders practical deployment in real-world applications. Although few-step generation methods such as MeanFlow substantially reduce the number of function evaluations (NFEs), two fundamental challenges still remain in extremely low-NFE models. The first is unreliability: in extremely low NFE environments, quality of generated samples becomes highly sensitive to the initial noise. The second is quality inconsistency: increasing the number of ODE steps does not guarantee monotonic improvements in general sample quality. We propose a hybrid noise skimming framework to overcome both issues through better initial noise selection in few-step generation. The proposed framework combines two complementary components: a learnable Noise Booster amortizes the initial noise selection at the distribution level by refining random noise toward more favorable regions and Single-Step Search which improves the sample quality as instance-wise adaptation using only a single additional search step. The learnable Noise Booster restores the reliability and raises the quality floor by refining initial noise and in contrast to increasing ODE steps, enlarging the search budget yields monotonic improvement in sample quality. Experiments at 44.1 kHz show that our framework achieves SOTA performance with only a one-step generation. The project page is available at: https://github.com/briskmuntjack/muntjack-20262027

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.