acceptodds
Under review as a conference paper at ICLR 2027

VMem-Bench: Attributing Long-Range Memory Failures in Causal Long-Video Generation

Abstract

Causal long-video generators increasingly treat memory as retrieval: to let characters and objects return after they leave the screen, they store past segments as frames, latents, or KV blocks and read back whatever best matches the current instruction or query. To test this bet, we propose **VMem-Bench**, a benchmark that scores memory both before and after generation and attributes each failure to the memory decision behind it. **Track-A** replays 92 real films (16,548 segments) and scores the references a memory hands to the generator before the next segment is revealed, with no generator involved. **Track-B** generates 30 long stories (2,824 segments) with one shared generator, run with retrieval memories and with two controls, no memory and each needed entity's previous appearance (*last-seen*), and scores whether required entities appear and look as before and whether removed ones stay out. Given the entities' names, retrieval finds most of the needed past, and its references keep returning entities closer to how they looked before at every gap. Its failures lie in decisions that similarity does not make: after the story changes an entity's look, it still hands over the old look, and the entity is drawn as described less often; it hands over more than is needed, costing precision; and it brings removed entities back, most of all when the prompt still names them. Last-seen references, which are told which entities are needed, gain comparable recall at lower precision and avoidance costs, yet they too hand over pre-change looks and fail when the instruction points to an earlier moment. Most inspected failures were settled before generation, in what memory stored or retrieved. Retrieval fails because it reads what resembles the instruction, not what is still true in the story; this suggests that it needs control that decides from the story when to read, what to read, and what to forget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.