acceptodds
Under review as a conference paper at ICLR 2027

PhyAudioBench: Benchmarking Physical Consistency in Audio Generation through Controlled Interventions

Abstract

Recent text-to-audio generation models have achieved substantial progress in perceptual quality and semantic alignment, yet existing evaluation largely focuses on whether generated audio sounds plausible, matches the prompt, and follows the requested temporal structure. Whether generated sounds behave consistently with their described physical conditions remains much less characterized. We introduce PhyAudioBench, a hierarchical benchmark for evaluating physical consistency in general audio generation. PhyAudioBench evaluates acoustic and prompt fidelity, temporal structure, and physical consistency across 240 structured prompts spanning eight acoustic scenes, three complexity levels, seven physical properties, and twelve controlled intervention types. Rather than judging physical realism from isolated waveforms, we construct minimally modified prompt pairs that alter a target physical condition and test whether the corresponding acoustic evidence changes in the expected direction. A shared multimodal observer grounds acoustic events in time, while property-specific signal evaluators perform physical measurement. Evaluating ten representative audio generators with fifteen metrics reveals that physical-consistency scores are consistently lower than lower-level scores, all ten models change rank between Level1 and Level3, and opposite-direction responses constitute a prominent failure mode under controlled interventions. These results expose a gap between generating audio that sounds plausible and generating audio that behaves consistently in physical world.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.