acceptodds
Under review as a conference paper at ICLR 2027

REEL: Learning to Retrieve, Evaluate, and Evolve Visual Memory for Any Long-Horizon Multi-Shot Video Generation

Abstract

Memory-based long-video generation aims to combine user control with long-range visual consistency. We organize this process into three functional components: a planner, an adapter, and a generator. Representative multi-shot systems couple memory to generator-specific interfaces or use predefined agent workflows, rather than independently training a memory policy for cross-shot objectives. We introduce REEL, a general-purpose visual-memory planner. Through a unified tool interface, REEL learns to retrieve references, create missing visual evidence, and update persistent memory. Backend-specific adapters translate its evidence plans into supported conditioning inputs, allowing the same policy to guide different black-box video generators. To train REEL, we construct two high-quality datasets, REEL-SFT-4k and REEL-RL-1k, covering diverse interaction patterns and prompt types. Training proceeds in three stages: supervised fine-tuning, single-shot Group Relative Policy Optimization (GRPO), and multi-shot GRPO. Multi-shot GRPO carries policy-generated memory across consecutive shots, so later outcomes guide earlier reference selection and memory updates. REEL-Bench tests identity preservation and intended state changes under text-only requests and timed reference-image uploads. On JoyAI-Echo 1.5 and MiniMax-H3, REEL improves cross-shot consistency on ST-Bench by 15.1% and 20.7%, and Core-100 on our proposed REEL-Bench by 2.41 and 8.32 points, respectively, over the no-memory baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.