ELTA: Event-Level Text-to-Audio Generation with Labels and Descriptions
Abstract
Sound design for films and games requires control over individual sounds, their acoustic details, and their timing. Coarse event labels leave those details unspecified, while requests involving several sounds also require each description to be associated with its intended event. We introduce ELTA, which uses explicit sound labels together with detailed natural-language descriptions within user-defined event intervals. ELTA encodes each sound label separately while retaining the full description of its event, then learns how these conditions contribute to one jointly generated scene. Multiple sounds can share an event, and separate events can overlap without merging their requests. For supervision, we curate and annotate 373,733 source audio clips with sound labels and detailed acoustic captions, then compose training scenes while preserving each recording's labels, description, and placement interval. Our results show improved sound coverage and description matching over temporal-generation baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.