acceptodds
Under review as a conference paper at ICLR 2027

CaReVOS: Causal Reasoning Across Missing Events for Video Object Segmentation

Abstract

Reasoning video object segmentation requires models to identify and segment objects from implicit queries that depend on understanding events and object relations. Existing settings, however, largely assume that the interactions defining these roles are observable. To study causal reasoning across missing events, we formulate a partial-observation setting in which models infer and ground the relevant participants from observations before and after an unobserved interaction. Based on this formulation, we introduce CaReVOS, a benchmark comprising 1,432 videos, 6,386 reasoning expressions, and 373,828 object mask annotations. It includes consequence, source, multi-solution, and counterfactual questions, with human-verified answers and masks. Because the interaction is hidden, answering these questions requires relating the observed setup to its outcome before selecting the target. To improve both this inference and the resulting segmentation, we develop Event-OPD. Supervised fine-tuning teaches target grounding, while privileged on-policy distillation lets a full-event teacher guide a student that sees only the opening and closing segments. Extensive evaluations show that existing referring and reasoning segmentation models struggle under this setting, while adaptation to CaReVOS yields substantial gains. Our distilled model achieves 58.33 , 61.98 , and 60.16 &, exceeding the strongest trained baseline by 2.21 points and its supervised initialization by 1.62 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.