acceptodds
Under review as a conference paper at ICLR 2027

Generative co-evolution: rethinking weakly-supervised video highlight detection in the MLLM era

Abstract

Weakly-supervised video highlight detection (WSVHD) aims to localize highlight moments using only video-level labels. While Multimodal Large Language Models (MLLMs) have strong reasoning capabilities, adapting them to WSVHD remains underexplored and is severely bottlenecked by sparse supervision and complex topic semantics. Conventional supervised fine-tuning (SFT) struggles under this signal scarcity, whereas semi-supervised approaches are constrained by static unlabeled datasets and laborious manual curation that stall rapid model iteration. To address this problem, we introduce a novel generative-adversarial post-training framework that drives MLLM performance improvement for WSVHD without human-curated unlabeled data. Rather than isolated multiple instance learning, we introduce a generative model that co-evolves with the MLLM using an efficient video synthesis tool. Through parameter updates guided by Group Relative Policy Optimization (GRPO), the generative model moves beyond a static data generator and synthesizes realistic, topic-relevant training video features to increase the MLLM's discriminative capability for WSVHD. Extensive experiments demonstrate that our co-evolutionary paradigm significantly outperforms standard supervised fine-tuning and classical semi-supervised learning, establishing new state-of-the-art performance for WSVHD across three public benchmarks. Code will be made publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.