Stage-adaptive Token Selection for Efficient Omni-modal LLMs
Abstract
The unified audio-visual understanding capability of omni-modal large language models (om-LLMs) comes at the high cost of processing tens of thousands of visual and audio tokens throughout the LLM. To reduce the cost, training-free methods for om-LLM token selection are being actively developed, operating either pre-LLM or inner-LLM. Recognizing that the non-textual token redundancy varies across stages and LLM layers, we propose in this paper a novel stage-adaptive token selection (SEATS) method that operates in both pre-LLM and inner-LLM stages. Given an overall token budget measured by the token retention rate (TRR), SEATS performs top-down token budget allocation in an input-dependent manner. Specifically, a larger TRR is used for pre-LLM token selection, whilst within the LLM, all tokens are retained in shallow layers, increasingly decreased TRRs are applied in middle layers, and the non-textual tokens are fully removed in late layers. Extensive experiments on three representative om-LLMs and six benchmarks verify the viability of the proposed method. By retaining only 10% of visual and audio tokens, SEATS preserves 94.9% of the full-token performance on Qwen2.5-Omni while achieving a FLOPs reduction and a prefill speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.