acceptodds
Under review as a conference paper at ICLR 2027

AudioRouter: Audio-Guided Visual Compression for Streaming Video Language Models

Abstract

Streaming video language models require efficient visual compression to preserve task-relevant evidence under causal constraints, where future observations are unavailable. Existing approaches mainly rely on visual signals for token selection, which can miss event-related information under aggressive compression. Meanwhile, audio-visual models typically inject audio representations into the LLM, requiring audio-capable architectures. We propose AudioRouter, an audio-guided visual compression framework based on a simple principle: audio determines where to look, while vision determines what the LLM sees. Instead of introducing audio tokens into the LLM, AudioRouter uses synchronized audio to condition learnable queries and generate a routing distribution over native visual tokens. The compressed representation is formulated as , where audio controls the routing weights while the aggregated values remain purely visual. This preserves the standard visual-token interface and enables causal streaming inference. Experiments on long-video and streaming benchmarks show that AudioRouter maintains performance close to the full 196-token representation with only 64 tokens per frame, retaining about one-third of the visual tokens while substantially reducing LLM inference cost. https://anonymous.4open.science/r/AudioRouter-1F96.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.