Generative co-evolution: rethinking weakly-supervised video highlight detection in the MLLM era
Abstract
Weakly-supervised video highlight detection (WSVHD) aims to localize highlight moments using only video-level labels. While Multimodal Large Language Models (MLLMs) have strong reasoning capabilities, adapting them to WSVHD remains underexplored and is severely bottlenecked by sparse supervision and complex topic semantics. Conventional supervised fine-tuning (SFT) struggles under this signal scarcity, whereas semi-supervised approaches are constrained by static unlabeled datasets and laborious manual curation that stall rapid model iteration. To address this problem, we introduce a novel generative-adversarial post-training framework that drives MLLM performance improvement for WSVHD without human-curated unlabeled data. Rather than isolated multiple instance learning, we introduce a generative model that co-evolves with the MLLM using an efficient video synthesis tool. Through parameter updates guided by Group Relative Policy Optimization (GRPO), the generative model moves beyond a static data generator and synthesizes realistic, topic-relevant training video features to increase the MLLM's discriminative capability for WSVHD. Extensive experiments demonstrate that our co-evolutionary paradigm significantly outperforms standard supervised fine-tuning and classical semi-supervised learning, establishing new state-of-the-art performance for WSVHD across three public benchmarks. Code will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.