JEC-Omni: Token Compression for Efficient Omnimodal Large Language Models via Joint Entropy Optimization
Abstract
Omnimodal Large Language Models (OmniLLMs) have demonstrated remarkable capabilities in audio-video understanding, but encoding continuous signals yields massive token sequences that impose enormous computational overhead and memory footprints. While training-free token compression offers a practical solution, existing omnimodal methods fail to precisely identify both intra- and cross-modal redundancies, struggling to stably maintain the full-token baseline accuracy under target token budgets. In this paper, we propose JEC-Omni, a novel training-free token compression framework that explicitly preserves high joint entropy for OmniLLMs. Our approach systematically eliminates redundancies through two core modules: MAST-P, which removes intra-frame spatial and inter-frame motion redundancies in the visual stream, and SICM-P, which leverages intra-modal temporal consistency and cross-modal semantic overlap to prune the audio stream. Extensive evaluations on the Qwen2.5-Omni and Qwen3-Omni architectures demonstrate that JEC-Omni achieves a superior trade-off between inference efficiency and accuracy. Remarkably, our method maintains 97% of the original model accuracy while retaining merely 25% of the raw audio-video tokens, reducing the end-to-end inference memory footprint by 1.1x to 1.8x and decreasing the Time-to-First-Token (TTFT) by 1.5x to 2.5x.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.