acceptodds
Under review as a conference paper at ICLR 2027

Atom-TTS: A lightweight on-device TTS system adopting cascaded prediction and frame-driven vocoder

Abstract

On-device speech synthesis needs to balance speech quality under a low computing budget. However, existing methods struggle to model high-quality intermediate audio representations with a small number of parameters. Moreover, high-quality vocoders incur substantial inference overhead on edge devices with constrained computing resources. To address these issues, we propose Atom-TTS. Starting from acoustic prediction and waveform reconstruction, it reduces inference costs while maintaining synthesis quality. First, we propose the Cascaded Parameter Acoustic Predictor (CPAP), which adopts a TCN encoder for cascaded encoding and establishes explicit conditional relationships among parameters to efficiently predict intermediate representations. It compresses trainable parameters to 161k without degrading prediction quality. Next, we propose the Frame-Driven WORLD vocoder (FD-WORLD). It constructs frame-level spectra using spectral envelope priors, harmonic excitation and noise excitation, and generates waveforms via short-window inverse Fourier transform and overlap-add, boosting the audio synthesis speed of the vocoder by about 5 times. In experiments, Atom-TTS achieves an RTF of 0.0065 on a single-core CPU, corresponding to 155 times faster than real time. In a subjective listening test with 15 listeners on a shared corpus, it reaches a MOS of 3.24, outperforming a same-corpus baseline with 3.2 times more parameters by 1.32 MOS and another with 9.7 times more parameters by 0.99 MOS. The backend replacement is not distinguishable under natural listening. 80.5% of trials are judged as having no audible difference. Atom-TTS thus achieves a balance among quality, model size and efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.