acceptodds
Under review as a conference paper at ICLR 2027

VibeVoice-ASR-Streaming: Who Spoke What in Real Time

Abstract

Real-time recognition of multi-speaker conversations is becoming increasingly important. Such systems need to transcribe speech, identify the speaker of each segment, and produce speaker-attributed transcripts with low delay. Existing end-to-end speaker-attributed ASR models jointly perform transcription and speaker attribution. However, they generally process complete recordings. Meanwhile, streaming LLM-based ASR has mainly focused on single-speaker transcription and generally lacks speaker-attribution capability. To address this gap, we present VibeVoice-ASR-Streaming, an end-to-end framework for streaming speaker-attributed ASR. It interleaves fixed-duration audio chunks with previously generated speaker labels and text in a single autoregressive context, which preserves speaker identities across chunks and enables speaker-attributed transcription in a single process without a separate diarization stage. On five recognition benchmarks, our 7B model achieves the lowest mean recognition error among the compared streaming systems. Across 13 speaker-attributed settings, it achieves the best or tied-best cpWER/cpCER in 12 settings. The expected latency is 2.00s. We we will release the 1.5B and 7B model weights together with the inference code.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.