NanoVoco: Upsampling-Aware Rank-Width Reallocation for Lightweight Vocoding
Abstract
Deploying text-to-speech (TTS) systems on resource-constrained edge devices requires lightweight neural vocoders. Under extreme compression, reducing channel width degrades synthesis quality. Although depthwise-separable convolutions (DSCs) allow wider layers at comparable parameter budgets, replacing standard convolutions (StdCs) with DSCs degrades quality even further in our experiments. We observe that this replacement involves reallocating parameters from convolutional rank to channel width, and find that the benefits of this rank-width reallocation depend strongly on the temporal upsampling rule. Replacing zero-insertion (ZI) with sample-repetition (SR) upsampling substantially improves the benefits of rank-width reallocation. We observe this interaction across different discriminator designs, generator backbones, and parameter budgets, with the strongest effect under extreme compression. Based on this observation, we propose NanoVoco, with only 72.3K parameters and UTMOS-measured synthesis quality comparable to the evaluated Parallel WaveGAN model and higher than FreGrad. Integrated into a complete TTS pipeline, its quantized implementation enables real-time streaming speech synthesis on a microcontroller using only on-chip memory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.