Prosodic Video Narration from Streaming Audio-Visual Events
Abstract
Compelling spoken narration requires not only the right words, but also prosody that reflects the significance of unfolding events. Existing streaming video models focus on what to say, while text-conditioned speech models lack the perceptual context needed to determine how to say it. We introduce Prosodic Video Narration, a task of continuously generating spoken descriptions whose content and prosody are grounded in streaming audio-visual events. To support this task, we construct a sports commentary dataset with aligned video, environmental audio, commentary text, and word-level prosodic annotations. We further propose ProsodyMLM, a streaming multimodal language model that models discrete word-level prosody from audio-visual context. Our design enables efficient autoregressive prosody prediction through lightweight adapters and prosody injection, while a two-stage training strategy transfers expressive speech capabilities before grounding them in multimodal events. Across quantitative metrics and a user study, ProsodyMLM generates narration that better aligns with reference speech and underlying events than existing baselines, improving pitch, speaking rate, and emotional alignment while maintaining competitive energy alignment, with particularly strong gains during highly expressive moments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.