GASA: A Structured Alignment Benchmark for Text-to-Audio Generation
Abstract
Text-to-audio generation has advanced faster than the methods used to evaluate it. Models are known to miss parts of a prompt, but prevailing measures judge a generated clip against its prompt as a whole and cannot show which parts are missed or where research should focus. We introduce the Generative Audio Structured Alignment (GASA) benchmark. GASA treats prompt fidelity as structured alignment, organizing the properties of a scene description into 25 fields and judging each field with a large audio-language model. Because such judges may answer from text-only priors without listening, we validate judge probes on real caption-audio pairs against a Gaussian-noise control, keeping for each field the probe whose accuracy depends most on the audio. Human raters answering the same probes on generated clips confirm that the judge detects 23 of the 25 fields in generated audio and set the ceiling each field score can attain. Using this validated judge, we evaluate 11 models on 2,500 generation prompts (100 per field). Scores vary far more across fields than across models, and every model leaves several fields near the control. Source identity is strongest, static acoustic properties weakest, and pitch and timbre are rendered more faithfully as changes than as static levels. GASA pinpoints where each model fails, giving researchers specific targets and a measure of progress toward models that realize every part of a prompt. https://gasa-demo.github.io/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.