NEOPrune: Non-rEconstructible Orthogonal Token Pruning for Omni Modal Models
Abstract
Omni-modal large language models (OmniLLMs) process large numbers of audio and video tokens, incurring substantial computational and memory costs. Existing token pruning methods primarily use token importance or pairwise similarity, leaving the collective replaceability of layer-wise token updates underexplored. We introduce NEOPrune, a training-free two-stage method that estimates replaceability by reconstructing orthogonalized layer-wise updates, focusing on information newly introduced at each layer. In the Pre-LLM stage, it combines within-modality reconstruction with encoder attention; in the Inner-LLM stage, it estimates video-update replaceability conditioned on audio and combines it with text-query attention. Across four audio–visual benchmarks and multiple OmniLLMs, NEOPrune generally achieves higher accuracy than evaluated pruning baselines at matched token-retention ratios, while retaining near-full-token performance with only 35–45% of audio–visual tokens.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.