acceptodds
Under review as a conference paper at ICLR 2027

WIRE: Weight Space Implicit Representation Encoding of Frozen Video LLMs

Abstract

Video language models represent a video as a long sequence of visual embeddings in a decoder's input. Every visual embedding is processed by every decoder layer during prefill, stored in the key value cache, and attended by every generated token, so a long video raises prefill, memory, and generation cost at once. We present WIRE (Weight Space Implicit Representation Encoding), which reduces this recurring cost by moving part of the video conditioned computation of a frozen video language model from its input activations into a temporary low rank update of its weights. WIRE keeps a short explicit visual context, so the update acts as a correction rather than a replacement of the video. WIRE first compresses the full resolution embeddings of each frame into a small set of carrier tokens using a question conditioned bidirectional selective scan, and physically contracts the decoder input so every decoder layer, cache entry, and decoding step operates on the shorter sequence. Hypernetworks at selected decoder layers read the compressed forward pass together with a readout of the full resolution features removed by contraction and generate LoRA updates to the attention and feedforward projections. The updates are merged after prefill, so autoregressive decoding runs the unmodified frozen backbone. On Qwen3-VL-8B, WIRE with 16 carrier tokens per frame contracts the visual stream by , reduces decoder prefill time by at 256 frames and at 1024 frames, shrinks the key value cache proportionally, and keeps per token decoding latency approximately constant. Across nine benchmark settings at 256 frames, it obtains the highest average accuracy among efficient video language models we evaluate under a common 256 frame budget. A paired update analysis toggles only the generated update while holding the compressed visual context fixed. Applying the generated update raises whole benchmark accuracy on Video-MME, MVBench, and LongVideoBench by up to points, with the largest gains on compression sensitive capabilities such as counting and action count, establishing that the generated update does task relevant work beyond the carrier tokens.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.