acceptodds
Under review as a conference paper at ICLR 2027

ScenA: Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

Abstract

Existing reference-conditioned multi-speaker systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings. These systems operate within speech-only pipelines that produce clean vocal sequences without the ambient texture of real conversations. Our method, , conditions a text-to-audio flow-matching foundation model, pretrained on large-scale in-the-wild data, directly on multiple reference voices and a free-form natural language prompt that describes an entire multi-speaker audio scene. Leveraging such a foundational model allows us to inherit its capacity for natural, non-studio audio: background noise, room acoustics, overlapping dialogue, and spontaneous paralinguistic events, while adding multi-speaker control without any per-turn structure. Concretely, reference latents are concatenated into the model's token sequence and distinguished by lightweight identity-aware positional encodings. However, we identify a critical obstacle to this approach: the . During training under standard noise schedules, the model can identify the matching reference by acoustic similarity to the noisy target, bypassing the text prompt entirely. We address this with a high-noise-biased timestep distribution that forces the model to rely on the text prompt for speaker assignment. We first evaluate broad scene-generation capabilities, demonstrating control over vocalizations, discrete events, emotional delivery, ambience, and simultaneous speech. Additionally, we evaluate on the CoVoMix2-Dialogue benchmark, where it outperforms specialized speech-only multi-speaker systems. Our results demonstrate the advantage of using a general-purpose audio model conditioned on a free-form scene description, rather than passing structured dialog scripts through a speech-only pipeline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.