DashengTokenizer: Acoustically Augmented Semantic Representations for Audio Understanding and Generation
Abstract
Audio understanding and generation place different demands on representations: accurate reconstruction does not ensure useful understanding features, while semantic features may not directly support sound synthesis. DashengTokenizer connects these capabilities by reusing continuous, high-dimensional features from a frozen audio encoder and injecting acoustics through a parallel lightweight Mel projection. The unquantized fusion preserves the original feature dimension. The projection and waveform decoder are jointly trained in one adaptation stage; understanding tasks read the fused features, while conditional generators predict the same type of representation before decoding. Experiments demonstrate competitive reconstruction across speech, music, and environmental sound and advantages over compared encoders and tokenizers on several sound and music understanding tasks. Full-sequence flow matching and continuous autoregression demonstrate use as a generation target. In multi-domain comparisons with the same DiT backbone configuration, the representation and decoder outperform an acoustic VAE scheme on text-to-audio, music, and speech metrics. Acoustic adaptation can thus extend pretrained understanding features into a continuous representation connecting reconstruction, understanding, and conditional generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.