acceptodds
Under review as a conference paper at ICLR 2027

Memory Internalization via Test-Time Training for Long-Video Understanding

Abstract

Long-video understanding often requires answering many questions about a single video whose evidence is distributed across extended temporal contexts. Conventional video-language models repeatedly encode frames or retrieve visual evidence for every question. Parameterized video memory offers a write-once alternative, but existing methods rely primarily on synthesized text and fixed adaptation configurations, leaving open where video knowledge should be written and how visual evidence should shape the memory. We introduce Memory Internalization via Test-Time Training (MIT), a per-video adaptation framework that writes each long video into compact, video-specific adaptive parameters for video-free, text-only question answering. Under compact multi-scale supervision, causal tracing measures the causal memory utility of each model layer for video-grounded answering and selects video-specific adaptation locations. Visual memory distillation then trains the text-only memory to reproduce the representation change induced by the relevant video clip, preserving the internal effect of visual evidence beyond textual supervision. Across LVBench, LongVideoBench, and Video-MME, MIT achieves data-efficient video-free question answering with substantially fewer synthesized samples than prior parameterized video-memory methods, demonstrating that effective video internalization depends on both where knowledge is written and how visual evidence is distilled into the model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.