acceptodds
Under review as a conference paper at ICLR 2027

EvoComp-2: Learning to Compress Audio-Video Tokens for Omnimodal Large Language Models via Topic-Guided Evolutionary Labeling

Abstract

Omnimodal Large Language Models (OmniLLMs) achieve strong performance on audio-video tasks but incur substantial computational overhead due to the large number of audio and video tokens. EvoComp proposes to adopt evolutionary search to construct token-retention supervision and train a lightweight compressor for visual token compression. In this work, we extend EvoComp to OmniLLMs and propose EvoComp-2 for efficient audio-video token compression. Extending EvoComp from static visual inputs to synchronized audio-video sequences introduces new challenges arising from temporal redundancy and query-dependent modality importance. To address them, EvoComp-2 introduces topic-guided temporal clustering that adapts a topic model to video frames and partitions audio-video sequences into semantically coherent intervals before evolutionary labeling, enabling more structured and efficient token selection. We further construct modality-level importance supervision to capture the relative contribution of audio and video to different queries. For compressor training, we introduce a temporal-window reweighting loss to emphasize information-rich temporal regions and a modality-level importance loss to learn query-dependent modality preference, together with the GHM loss inherited from EvoComp. Extensive experiments demonstrate that EvoComp-2 achieves favorable accuracy-efficiency trade-offs. On WorldSense, EvoComp-2 preserves 98.7% of the original accuracy while retaining only 35% of the audio-video tokens. In terms of inference efficiency, it achieves a 1.6x end-to-end speedup on DailyOmni with 3.3x token compression.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.