acceptodds
Under review as a conference paper at ICLR 2027

Beyond RGB Reconstruction: Predictive State-to-Token Adaptation for Video MLLMs

Abstract

Recent advances in multimodal large language models (MLLMs) have improved video understanding, motivating more efficient visual interfaces. Conventional pipelines reconstruct compressed video into RGB frames and re-encode them, despite predictive decoders already producing temporally conditioned states. We introduce **PreSTA** (Predictive State-to-Token Adaptation), a receiver-side interface that maps these states into the intermediate visual space of a frozen MLLM. Under the same eight-group output budget as sparse RGB, PreSTA aggregates 32 P-frame states into compact visual groups, bypassing RGB reconstruction and early visual encoding for selected observations while preserving predictive state updates. A compact RGB teacher trained with privileged source-frame supervision guides the codec-domain adapter through feature- and semantic-level alignment; the codec, bitstream, and MLLM remain unchanged. Across four codec operating points on Video-MME and MVBench with Qwen3-VL-8B and Qwen3-VL-4B, PreSTA stays within 0.3 points of a stronger 32-frame RGB-Fusion baseline on both benchmarks and model scales, and improves Video-MME over the standard sparse Reconstructed RGB-8 pipeline by an average of 6.2 and 6.1 percentage points. Relative to RGB-Fusion, it achieves this with 49.2% fewer visual-front-end weights, 88.5% fewer MACs, and 93.4% lower measured latency from decoder outputs to visual block 2.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.