acceptodds
Under review as a conference paper at ICLR 2027

AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation

Abstract

Text-to-audio (TTA) generation has made substantial progress in synthesizing realistic audio from natural-language descriptions. However, determining whether generated audio faithfully satisfies complex instructions remains challenging. Widely used global text–audio similarity metrics provide limited insight into which specific semantic requirements are satisfied or omitted. We introduce AudioScape-TTA, a structured and complexity-aware benchmark for fine-grained TTA evaluation. The benchmark represents soundscapes through verified annotations of scene context, sound effects, background music, and speech, and characterizes sample complexity using annotation-derived event density and audio-component structure. From these annotations, we derive fixed semantic rubrics covering event presence, acoustic attributes, and specified speech content. Event and attribute requirements are evaluated by an audio-language model, while speech content is assessed through a separate ASR-based coverage pathway. AudioScape-TTA contains 2,258 audio–text pairs and 25,707 semantic rubrics, enabling scalable and interpretable analysis of TTA systems. Evaluation of 13 open-source models reveals distinct profiles in event realization, attribute control, speech-content coverage, and compositional soundscape coverage. On a human-study subset, the proposed scores achieve higher model-level rank correlation with human semantic judgments than CLAP similarity, supporting structured semantic assessment beyond global text–audio alignment. Anonymous project page:https://anonymous.4open.science/w/AudioScape-T2A-426A/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.