acceptodds
Under review as a conference paper at ICLR 2027

Specialize Before Fusion: Modality-Specific Routing for Multimodal Video Highlight Detection

Abstract

Highlight detection and video summarization require identifying salient moments from heterogeneous videos where visual and acoustic evidence can vary over time. Existing multimodal methods mainly focus on how modalities are fused, often before adequately modeling their distinct temporal characteristics. We introduce a specialize-before-fusion framework that first learns independently parameterized visual and audio temporal representations and then adaptively combines them using time-varying routing weights, optionally conditioned on a semantic query. Across four benchmarks spanning long-form sports broadcasts and web video, our method achieves state-of-the-art performance, improving over the strongest baseline by up to 6.9 mAP on long-form sports highlight detection while incurring lower inference cost. Ablations consistently favor modality-specific over shared temporal modeling and adaptive over fixed fusion, and representation analyses show visual and audio branches become more distinct specifically near salient moments. These results support specializing heterogeneous modalities before multimodal integration.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.