Binding Voices to Characters: Disentangled Cross-Modal Alignment for Multi-Speaker PercepTion
Abstract
Humans effortlessly perceive crowded multi-speaker scenes, resolving *who* is speaking, *when*, *where*, and *what* they say in real time. Replicating this capability in machines, however, remains a fundamental open challenge rooted in cross-modal speaker ambiguity that existing end-to-end models fail to resolve. A natural remedy is to scale training through data engineering. However, scaling training data alone may not sufficiently constrain the correspondence between speech segments and visible speakers, motivating explicit structural supervision for speaker binding. To address these limitations, we propose **CAST** (disentangled **C**ross-modal **A**lignment for multi-**S**peaker percep**T**ion), a framework that endows model with genuine multi-speaker perception capability through two complementary structural interventions. First, Speaker-aware Audio Encoding (SAE) restructures the audio encoder to explicitly segment the audio stream by speaker identity. Second, an Audio-Conditioned Spatial Modulation (ACSM) module performs audio-visual contrastive alignment to generate a Speaker Heatmap Prior, whose spatial tokens are superimposed onto the visual tokens to produce speaker-conditioned visual representations. CAST is then trained end-to-end via supervised fine-tuning followed by reinforcement learning to align perception with structured reasoning. We further release **MSR-QA** (Multi-Speaker Reasoning QA), a large-scale multilingual dataset comprising over **8000+ hours** of effective speech video and **11 million** QA pairs across movies, TV dramas, variety shows, and interview programs. Extensive experiments demonstrate that CAST significantly outperforms existing approaches by up to **34.30** percentage points on cross-speaker temporal and spatial reasoning, while providing transparent and verifiable reasoning processes. **Web Demo:** [https://cast-iclr.github.io/CAST](https://cast-iclr.github.io/CAST)
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.