acceptodds
Under review as a conference paper at ICLR 2027

Long-Term Visual Understanding with Latent Token Memory

Abstract

While Large Multimodal Models (LMMs) have demonstrated impressive capabilities in reasoning over static images, they face severe scalability bottlenecks when processing extensive visual streams, such as long-form videos or multi-page documents. Prevailing approaches typically resort to raw token concatenation, which incurs quadratic computational costs, or aggressive stateless pruning, which inevitably sacrifices fine-grained semantic details. To address this, we introduce MemVLM, a novel framework that treats visual history as an online information processing problem. Instead of buffering an ever-growing queue of tokens, MemVLM maintains a fixed-size latent belief state, acting as a sufficient statistic for the visual history. Our architecture employs a dual-head projection to decouple immediate visual grounding from long-term memory updates, utilizing a Gated Recurrent Memory Update (GRMU) mechanism to distill high-entropy evidence into the latent state. Guided by a "Minimal Write Principle," the model learns to selectively update memory only when necessary, effectively transforming the transformer into a state-space model with memory complexity. Empirical evaluations on comprehensive long-context benchmarks demonstrate that MemVLM significantly outperforms state-of-the-art token pruning methods, enabling the processing of thousands of frames without "lost-in-the-middle" degradation or memory explosion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.