acceptodds
Under review as a conference paper at ICLR 2027

RABBIT: Dynamic Budget Allocation Token Compression for Omni-modal LLMs

Abstract

Omnimodal large models integrate video, audio, and text modalities and have shown strong cross modal capabilities for real world interactive applications. However, long audio and video inputs introduce a large number of tokens, making inference expensive and hindering the deployment of omni modal large models in practical applications. Existing compression methods mainly rely on heuristic token importance metrics and fixed audio and video compression ratios, which often fail under aggressive compression and overlook sample level modality differences. We propose a reinforcement learning based audio and video token compression method RABBIT (Reinforcement Learning for Adaptive Budget Balancing in Token Compression) that learns a unified ranking of task relevant audio and video tokens for each sample. Our method uses modality specific scoring modules and joint audio and video rollout sampling to adaptively allocate the compression budget between audio and video while preserving full token model behavior through teacher logit distillation. Experiments on multiple audio and video understanding benchmarks show that our method outperforms the evaluated baselines under different compression ratios. On WorldSense, it preserves of full token performance with only audio and video tokens and outperforms DivPrune and SEATS by and percentage points, respectively. Further analysis shows that the learned sample level audio and video budget allocation is effective, improving SEATS by points on WorldSense and points on VideoMME at the retain ratio.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.