Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos
Abstract
Spatio-temporal scene graphs provide a principled representation for modeling evolving object interactions, yet existing methods remain fundamentally frame-centric: they operate primarily in 2D, reason about currently visible objects, and discard entities upon occlusion. We introduce World Scene Graph Generation (WSGG), which constructs a temporally persistent world scene graph encompassing observed and unobserved interacting objects. To support this task, we introduce ActionGenome4D, upgrading ActionGenome videos to 4D scenes. We study the task in two settings: supervised WSGG and zero-shot WSGG. For the supervised WSGG setting, we begin by adapting existing VidSGG methods to WSGG to establish strong baselines. Then, we propose a suite of methods (WorldWise, WorldWise+, WorldWise++) that captures joint visual and 3D object representations and treats occlusion as a natural masking signal, reconstructing missing features via camera-pose-conditioned cross-attention over an object's visible history. For zero-shot WSGG, we evaluate open-source MLLMs under strong context-based prompting (including our proposed graph retrieval framework UWSGG-GraphRAG) strategies with explicit geometric inputs and outputs. Extensive experiments demonstrate the efficacy of the proposed methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.