Learning to Reason with Persistent Object States for Video Instance Segmentation
Abstract
Video segmentation models preserve object identities by carrying information about each instance across frames. They typically update this history as new observations arrive. Under prolonged occlusion, reappearance, or interactions between similar instances, however, an incorrect update can overwrite a valid history and turn a local mismatch into persistent identity drift. This raises a question that current pipelines largely leave implicit: when should an observation be allowed to change an object's state? We introduce **POSReasoner**, a trainable, plug-and-play framework that makes this decision explicit by reasoning over persistent object states. For each object, POSReasoner carries a compact state that records identity, geometry, visibility history, and confidence. A sparse interaction graph links these states to candidate observations. Over several reasoning steps, POSReasoner compares each candidate with the object's history, resolves conflicting associations, and determines whether the state should be retained, updated, or reactivated. Intermediate supervision guides association and presence prediction, while a learned gate blocks unreliable updates. The resulting state is carried forward, so each decision affects subsequent association and segmentation rather than merely reranking a fixed set of predictions. POSReasoner is trained with standard video annotations while the base model remains frozen. This lightweight design allows it to augment diverse VOS and VIS architectures. Across long-term VOS and VIS benchmarks, POSReasoner consistently improves strong baselines, with the largest gains under occlusion and object reappearance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.