acceptodds
Under review as a conference paper at ICLR 2027

Beyond Uniform Pooling: Typed Object and Relation Evidence for Frozen Video LLMs

Abstract

Video language models spend a fixed visual-token budget on uniformly pooled patches or learned latents, so the tokens a model reads carry no record of which object, moment, or relation they describe. We present StateWeaver, which replaces part of the uniformly pooled grid with typed evidence that a frozen language model reads within the same token budget. Each video is parsed once, with frozen trackers and a frozen image encoder, into a reusable EvidenceBank of scene features, tracked objects with state timelines, and pairwise relations. For each question, an answer-free retriever takes the dense connector’s own 384-token sequence, keeps 256 positions as question-independent context, and writes native scene, object, and relation-selected object features onto the other 128; only the connector is trained. On the official NExT-QA test, StateWeaver reaches 58.1%, 15.0 points above the 384-token dense connector and 8.7 points above a resampler that reads the same memory and question, and it also improves three further video benchmarks. A factorial ablation finds positive main effects for object evidence, relation-selected object appearance, and the posterior teacher; beyond that teacher, conditioning retrieval on the question adds no separable benefit. All comparisons use one frozen 8B backbone at one 384-token budget, and the resampler contrast changes the layout, the write, and the training objective together.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.