acceptodds
Under review as a conference paper at ICLR 2027

Learning to Organize Continuous Audio Representations for Streaming Generation

Abstract

Streaming speech generation poses a representation challenge: acoustic latents must be compact and easy to predict incrementally, yet retain the detail needed to reconstruct natural speech and music. Systems that use an external semantic encoder must maintain a relationship between acoustic and semantic representations during generation. We propose a low-frame-rate continuous audio representation that incorporates semantic structure directly into reconstructable acoustic latents. The same representation serves as a generation target and supports autoregressive audio-history modeling, without requiring an external semantic encoder at inference. Block-causal processing makes the representation suitable for streaming, while soft semantic alignment provides a content prior without removing acoustic flexibility. We further organize channels into expandable ordered prefixes, allowing a single audio encoder–decoder to provide representations at different capacities and enabling the generator to adjust its learning priorities at a fixed capacity. Experiments on Seed-TTS Chinese speech generation and speech and music reconstruction show that this design offers a unified representation for streaming generation with semantic guidance, acoustic fidelity, and adjustable capacity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.