acceptodds
Under review as a conference paper at ICLR 2027

MOSS-Audio: Unified Audio Understanding Across Semantics, Acoustics, and Time

Abstract

Unified audio understanding requires integrating semantic content with fine-grained acoustic cues and temporal structure. We present MOSS-Audio, a family of audio-language models for unified understanding and reasoning across speech, environmental sound, and music. An audio encoder trained from scratch and DeepStack cross-layer feature injection provide the language decoder with acoustic features at multiple abstraction levels. Explicit time markers anchor features to elapsed time for timestamped transcription and time-aware question answering through autoregressive generation. Our annotation pipeline segments recordings at event boundaries and merges specialist outputs into unified captions that preserve acoustic detail and event chronology, while retaining intermediate annotations for task-specific instruction tuning. Large-scale pretraining and staged post-training yield instruction-following (Instruct) and reasoning (Thinking) variants. Across evaluation, MOSS-Audio achieves the highest average accuracy among evaluated open-source systems, including 30B-scale models. MOSS-Audio also leads evaluated open and proprietary systems in average speech-captioning score and achieves more accurate timestamp alignment than evaluated omni-model baselines. We will release model weights, training data, code, and complete training recipes to support further research in unified audio understanding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.