SCENORBIT: Grounding Scene Programs for Consistent Multiview Video Generation
Abstract
A controllable world generator must turn a requested scene into mutually consistent observations across cameras and time. Current systems distribute this responsibility across language interfaces, geometric controllers, and visual generators, leaving the agreement between them to emerge indirectly. We introduce SCENORBIT, which treats a structured scene program as a common reference for these components. A language agent specifies the scene, a geometric renderer grounds its entities in each camera, and a visual critic supervises their agreement. Shared compression and positional binding connect the rendered controls to the video representation before generation. This design supports precise scene editing while preserving the relationships between views. On nuScenes, the reported evaluation achieves 0.132 epipolar photometric error and 21.55 BEV detection mAP, with 32-frame, six-camera generation in 9.6 seconds. Training a detector on 20,000 generated clips yields 44.0 NDS, while combining real and generated data reaches 45.7 NDS. Component studies connect these gains to geometric grounding, aligned conditioning, and critic feedback.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.