ReVoC: High-Fidelity and Efficient Audio Generation via Restoration Bridges and GANs
Abstract
Although reconstructing high-fidelity audio from compressed representations is an underdetermined inverse problem, the corresponding inverse transforms or decoders can recover useful acoustic structure at low computational cost. In this paper, we propose ReVoC, a few-step generative framework that combines restoration bridges with adversarial training to refine these approximate reconstructions. Starting from the outputs of Mel inversion or neural codec decoding, ReVoC learns to restore clean audio in the complex STFT domain using an efficient Vocos-style 1D backbone. To improve perceptual quality across inference budgets, we introduce multi-NFE adversarial fine-tuning, where a step-conditioned discriminator guides the joint optimization of complete one-, two-, and four-step inference trajectories. A single generator per task thus supports all three budgets with quality comparable to models fine-tuned separately for each step count. On Mel-conditioned vocoding, ReVoC-Small uses the same eight-block backbone as Vocos and improves one-step PESQ from 3.592 to 4.142. With four-step inference, ReVoC-Base achieves a PESQ of 4.433 and approximately four times the GPU throughput of four-step BridgeVoC. ReVoC also consistently improves reconstruction quality in low-bitrate neural codec restoration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.