acceptodds
Under review as a conference paper at ICLR 2027

Empowering Unified Audio Language Modeling with Hierarchical Semantic Representations

Abstract

We present SL-Audio, an audio-language modeling framework that supports audio captioning, text-to-audio generation, and text-guided audio separation through hierarchical semantic representations. Unlike existing approaches that typically address these capabilities independently or rely on a single semantic granularity, SL-Audio organizes audio into multiple semantic layers, ranging from high-level semantic concepts to progressively finer acoustic representations. This hierarchical design enables different tasks to flexibly access representations at appropriate levels of abstraction, where higher-level representations provide semantic guidance and finer representations capture progressively richer acoustic details. Building on this hierarchy, SL-Audio enables a single LLM backbone to support audio captioning, text-to-audio generation, and text-guided audio separation within a unified framework, allowing different tasks to leverage different levels of the hierarchy according to their representation requirements. Experiments on these representative tasks demonstrate that the proposed hierarchical representations effectively support diverse audio-language tasks within a unified framework, providing an initial step toward unified audio-language modeling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.