acceptodds
Under review as a conference paper at ICLR 2027

CASTSCRIPT: ALIGNING “WHO” AND “WHAT” IN MULTILINGUAL AUDIO-VISUAL CAPTIONING

Abstract

Speaker-attributed transcription, i.e., jointly determining who speaks and what is said, is a pivotal yet under-explored capability in audio-visual captioning, held back by two gaps. First, existing dialogue-centric benchmarks are largely limited to English and Chinese and lack scene and acoustic annotations for difficulty-stratified diagnosis. Second, speaker attribution and utterance transcription are heterogeneous competencies, yet existing methods optimize them under a single aggregate objective, yielding supervision that is neither dense nor dimension-specific. To close these gaps, we introduce CastScriptBench and CastScript. CastScriptBench comprises 508 in-the-wild conversational videos over 13 languages, spanning 1,356 shots, 4,236 utterances, and 1,218 speakers, annotated with 8 scene and acoustic attributes for difficulty-stratified analysis and evaluated by a verifiable protocol that combines edit-distance-based transcription evaluation with checklist-based speaker attribution. CastScript is the first to adapt multi-teacher on-policy distillation (MOPD) to captioning: unlike conventional MOPD that assigns experts by task, we route by output dimension within a single sequence, supervising tag tokens with an attribution expert and text tokens with a transcription expert, replacing the sequence-level scalar reward with per-token distribution matching. On CastScriptBench, both proprietary and open-source models mis-attribute and mistranscribe, worsening sharply as speaker count and overlap rise. CastScript surpasses even the transcription expert in fidelity, attains the best open-source performance, and establishes MOPD as an effective recipe for multi-dimensional captioning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.