acceptodds
Under review as a conference paper at ICLR 2027

Dynamic Audio-Visual Correlation-Aware Token Compression for Omnimodal Large Language Models

Abstract

Omni-modal large language models (OmniLLMs) support unified audio-visual understanding. However, processing large numbers of audio and video tokens causes substantial computational and memory costs, making efficient token compression essential. Audio and video can provide unique or repeated information at different moments, while existing methods typically apply one fixed strategy throughout a sequence. Guiding one modality with the other may discard information unique to each modality when they are complementary, while compressing them separately may preserve redundant information when they overlap. Such fixed strategies fail to adapt to this changing correlation. To address this mismatch, we propose OmniTrim, a training-free framework for adaptive audio-visual token compression. Specifically, Information-Aware Budgeting allocates the token budget across temporal chunks according to their audio-visual activity and temporal variation. Correlation-Aware Selection then adaptively chooses between separate selection and joint token merging based on the local audio-visual correlation, preserving complementary information and removing redundancy. Experiments on six audio-visual benchmarks demonstrate that OmniTrim outperforms existing OmniLLM compression baselines. At 25% token retention, it preserves 98.8% of the full-token performance on Qwen2.5-Omni-7B, while delivering a 4.35× prefill speedup, increasing decoding throughput by 55.3%, and maintaining the lowest peak GPU memory usage and end-to-end latency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.