acceptodds
Under review as a conference paper at ICLR 2027

Scaling Unified Multimodal Models to Long-Horizon Interleaved Generation

Abstract

Unified multimodal models (UMMs) provide an interface for interleaved text-image generation, where generated visual states can directly inform subsequent reasoning and synthesis. However, existing systems are still primarily demonstrated on short trajectories or simplified visual content. We develop a large UMM natively trained on interleaved multimodal data, with strong capabilities in image synthesis, text rendering, complex composition, and multimodal planning, and study how to extend these native capabilities to substantially longer autoregressive trajectories. The main challenge is that generated images continuously accumulate in context: irrelevant or imperfect visual states interfere with future synthesis, while previously established characters and scenes must still be recovered after long intervals. We introduce a visual-memory framework that maintains persistent visual references, selectively retrieves scene-relevant historical states while preserving the complete semantic history, and uses few-step self-rollout to adapt the model to imperfect generated histories. We further introduce LongMIG-Bench for evaluating long-horizon text-to-multi-image and image-to-multi-image generation. Our model achieves high-fidelity long-horizon interleaved generation with strong character and scene consistency, demonstrating how visual memory can scale native interleaved generation to substantially longer trajectories.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.