acceptodds
Under review as a conference paper at ICLR 2027

Remembering What Matters: Balanced Memory for Ultra-Long Streaming Videos

Abstract

Multimodal Large Language Models (MLLMs) for streaming videos compress an unbounded stream of past frames into a fixed-size memory to stay within a limited LLM context budget. However, existing methods often produce unbalanced memory representations: heuristic writes induce a recency bias that coarsens older content, while attention-based writes concentrate on a few dominant events; both losing distant evidence as streams grow longer. Thus, we reformulate streaming memory compression as a **balanced temporal coverage** problem and propose **StreamBalance**, whose fixed-size memory is fully learnable and trained end-to-end through two mechanisms. A *state-aware, entropy-regularized optimal transport* formulation that routes tokens to their best-fit memory slots under a capacity constraint that keeps utilization even on write, and *adaptive decay-rate-aware retention mechanism* in which each slot learn its own persistence: fast slots merging recurring redundancy, and slow slots preserving rare events; so answer-bearing evidence is neither crowded out on write nor decayed away over time. Together, these mechanisms let StreamBalance set a new state of the art on two streaming benchmarks, such as % overall gain on OVOBench. On two ultra-long benchmarks spanning up to hours, it maintains nearly stable accuracy where prior methods collapse, outperforming prior streaming model by % on EgoLife-QA at lesser compute and lower latency. Despite being trained for streaming, it further generalizes competitively to three offline long-video benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.