Execution Before Memory: Diagnosing Memory Failures in Video World Models
Abstract
A memory test for a video world model should say whether the scene was set, whether the asked change was carried out, and whether the last frame shows the result. Current tests fold these into one number, so a low score is easy to read as forgetting. When the change is not carried out, the hidden state transition cannot appear, so the test is not really a test of memory. We introduce MIST, an evaluation that separates the change from the memory it should leave behind. The model first sets the scene while the change stays hidden, and only later is asked to show it, with leakage before that point kept under control. This separation shows that the present bottleneck is following the instruction for that execution. The same gap appears in general memory evaluations. On that basis we propose M3-Agent (Make the Missing Move), which recursively revises the details of the input instruction so that the model successfully carries out the requested instruction. Relative to the baseline method, the change is shown more than twice as often. On the benchmarks, WRBench's score T rises by 26.1% and MBench-T's score U by 52.7%. The method thus meets the problem in memory evaluation from the execution side: it releases the model's potential, and it builds a stricter comparison.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.