RABBIT: Dynamic Budget Allocation Token Compression for Omni-modal LLMs
Abstract
Omnimodal large models integrate video, audio, and text modalities and have shown strong cross modal capabilities for real world interactive applications. However, long audio and video inputs introduce a large number of tokens, making inference expensive and hindering the deployment of omni modal large models in practical applications. Existing compression methods mainly rely on heuristic token importance metrics and fixed audio and video compression ratios, which often fail under aggressive compression and overlook sample level modality differences. We propose a reinforcement learning based audio and video token compression method RABBIT (Reinforcement Learning for Adaptive Budget Balancing in Token Compression) that learns a unified ranking of task relevant audio and video tokens for each sample. Our method uses modality specific scoring modules and joint audio and video rollout sampling to adaptively allocate the compression budget between audio and video while preserving full token model behavior through teacher logit distillation. Experiments on multiple audio and video understanding benchmarks show that our method outperforms the evaluated baselines under different compression ratios. On WorldSense, it preserves of full token performance with only audio and video tokens and outperforms DivPrune and SEATS by and percentage points, respectively. Further analysis shows that the learned sample level audio and video budget allocation is effective, improving SEATS by points on WorldSense and points on VideoMME at the retain ratio.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.