acceptodds
Under review as a conference paper at ICLR 2027

OmniPartition: Query-Guided Region Refinement for Audio–Visual Token Pruning

Abstract

Omnimodal large language models (OmniLLMs) jointly process audio, video, and text through a shared decoder sequence, but the resulting long multimodal context makes prefill increasingly expensive. Training-free token reduction offers an effective solution, yet existing methods often overlook the spatial and temporal structure of multimodal inputs or rely on fixed modality allocations and auxiliary models. We introduce OmniPartition, a training-free framework for query-guided, structure-aware audio–visual token pruning. OmniPartition first performs pruning after the first decoder block, where contextualized representations provide a reliable query-affinity signal. It then partitions video and audio tokens along their native spatial and temporal coordinates and adaptively refines these regions according to their query-weighted representation error. A shared integer budget is allocated through the marginal gain of each refinement, allowing the method to jointly determine which regions to preserve and how many tokens to retain without predefined modality ratios or temporal windows. The resulting compact sequence can be directly processed by standard FlashAttention. Across six audio–visual benchmarks and multiple OmniLLM backbones, OmniPartition consistently preserves performance under substantial token reduction. On Qwen2.5-Omni-7B, it retains 95.27% of full-token performance at 20% retention while achieving a 3.34 prefill speedup.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.