MEMOIR: Region-Level Graph Memory for Long-Term Video Object Segmentation
Abstract
Long-term video object segmentation must preserve target identity across prolonged absence, severe occlusion, and similar distractors, a requirement that ties each prediction to historical appearance evidence far beyond the recent observations. Existing memory designs retain this history as whole frames within a bounded window and expose all of it to every prediction. Informative partial appearances therefore vanish with the discarded frames, and every prediction attends to the entire store regardless of what it requires. To address this limitation, we propose Memory with Evidence Modeled On Interconnected Regions (MEMOIR), a graph memory that stores, maintains, and exposes historical evidence at the level of spatially coherent regions. MEMOIR organizes the region-level memories of a frozen SAM-style base model into a graph whose edges encode spatial, similarity, and temporal relations. Retrieval exposes long-term evidence only where recent observations fail to support the prediction, and management maintains the graph within explicit node and token budgets according to temporal support and retrieval utility rather than age. The base model itself remains unchanged, and no training is required. Across seven video object segmentation benchmarks, MEMOIR improves on the SAM 3 base model with the largest gains on LVOSv2 and MOSEv2, whose long and crowded videos stress memory most, and exceeds SAM 3 on all four reappearance splits, where historical evidence determines the prediction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.