acceptodds
Under review as a conference paper at ICLR 2027

ARMem: Autoregressive Latent Memory for Streaming Video Generation

Abstract

Long-horizon autoregressive (AR) video generation requires compact memory that preserves persistent information while adapting to evolving content. Existing methods typically preserve history through post-hoc selection or compression of previously generated visual representations and reuse them as static context. While effective for preserving past information, these static memories cannot adapt to evolving video dynamics and may lead to quality degradation during long rollouts. In this work, we introduce **ARMem**, an autoregressive memory paradigm that extends AR modeling from video generation to memory evolution. Specifically, memory is represented as compact latent tokens that are jointly denoised with video tokens in a shared diffusion transformer. At each step, the previous memory state guides video synthesis, while newly generated content is incorporated into the next memory state to support subsequent generation. To mitigate optimization conflicts between video generation and memory evolution, we introduce Dual-LoRA Self-Attention, which decouples their parameter adaptations. We further introduce a Historical Readout Task that explicitly supervises the evolving memory to retain future-useful information, such as subject identity. Extensive experiments on LTX and Wan demonstrate improved short-video quality and long-horizon performance under bounded memory constraints.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.