VibeVoice-ASR-Streaming: Who Spoke What in Real Time
Abstract
Real-time recognition of multi-speaker conversations is becoming increasingly important. Such systems need to transcribe speech, identify the speaker of each segment, and produce speaker-attributed transcripts with low delay. Existing end-to-end speaker-attributed ASR models jointly perform transcription and speaker attribution. However, they generally process complete recordings. Meanwhile, streaming LLM-based ASR has mainly focused on single-speaker transcription and generally lacks speaker-attribution capability. To address this gap, we present VibeVoice-ASR-Streaming, an end-to-end framework for streaming speaker-attributed ASR. It interleaves fixed-duration audio chunks with previously generated speaker labels and text in a single autoregressive context, which preserves speaker identities across chunks and enables speaker-attributed transcription in a single process without a separate diarization stage. On five recognition benchmarks, our 7B model achieves the lowest mean recognition error among the compared streaming systems. Across 13 speaker-attributed settings, it achieves the best or tied-best cpWER/cpCER in 12 settings. The expected latency is 2.00s. We we will release the 1.5B and 7B model weights together with the inference code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.