acceptodds
Under review as a conference paper at ICLR 2027

OmniRelay: Information-Flow-Guided Token Compression for Omni-Modal LLMs

Abstract

Omni-modal large language models (Omni-LLMs) achieve strong performance on audiovisual understanding tasks, but processing long audiovisual token sequences incurs substantial computational and memory costs. To reduce these costs, existing token compression methods exploit local redundancy at the pre-LLM stage and task relevance at the within-LLM stage, but make limited use of global modality dependencies. To determine when these dependencies should guide compression, we analyze attention patterns and perform attention knockout across layers. We observe a depth-varying relay: audiovisual information is integrated through intra-modal interactions within and across windows in early layers, transferred to text in middle layers, and finally relayed to the final input token for prediction. We therefore propose OmniRelay, a training-free pre&within-LLM framework that adapts compression guidance across layers. Its Depth-Adaptive Relay Compression (DARC) progressively shifts from intra-modal to text guidance before removing the remaining audiovisual tokens at late layers. Complementing DARC, our proposed Saliency–Diversity and Coverage-Aware Compression (SDCC) reduces local redundancy before the LLM while preserving salient and diverse tokens with broad cross-window coverage. We evaluate OmniRelay on 3 Omni-LLMs and 5 audiovisual benchmarks. It retains 99.51% of full-token performance at 25% layer-averaged audiovisual token retention, with 79% fewer FLOPs and a 3.3× prefill speedup.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.