acceptodds
Under review as a conference paper at ICLR 2027

MoCo: Modality-Aware Decoder-Side Token Compression for Video–Audio OmniLLMs

Abstract

Long audio–video inputs make inference costly for omnimodal large language models (OmniLLMs), motivating token compression that preserves question-relevant evidence. Question-conditioned attention offers a natural signal for selecting which tokens to retain. However, we find that shallow question-conditioned attention favors audio even for visual questions, causing direct cross-modal ranking to allocate less capacity to visual evidence. Allocating audio and video budgets separately improves accuracy without changing token scores or within-modality rankings. Based on these findings, we propose MoCo, a training-free framework for decoder-side token compression that separates modality-budget allocation from question-conditioned token selection. It assigns audio and video shares of a joint token budget, then uses question-conditioned attention to rank tokens within each modality. Within a single prefill pass, MoCo prunes audio early to reduce computation and delays video selection to exploit more informative attention scores at later shallow layers. Across four OmniLLMs, MoCo achieves higher average performance than OmniZip at all three evaluated retention ratios. MoCo preserves 98.7% and 97.3% of full-token performance on Qwen2.5-Omni-7B and 3B, respectively, while retaining only 25% of multimodal tokens. At matched FLOPs, MoCo preserves up to 3.1 percentage points more of full-token performance than OmniZip on Qwen2.5-Omni-7B.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.