acceptodds
Under review as a conference paper at ICLR 2027

CoMO4D: Temporally Consistent Multi-Object 4D Generation from a Single Monocular Video

Abstract

Generating temporally consistent multi-object 4D data from a monocular video is fundamental to dynamic scene modeling, yet existing methods often struggle to retain semantic object organization and maintain spatial and temporal consistency under occlusion. We propose CoMO4D, a structured 4D generation framework that explicitly represents each persistent semantic object with a single shared canonical geometry and a sequence of frame-wise spatial states, thereby describing both the stable object structure and its time-varying spatial evolution throughout the video. CoMO4D combines monocular depth estimates with semantic cues to construct object-level point-cloud sequences. At its core is a spatiotemporal completion mechanism that integrates semantic-region and inter-object relational reasoning, cross-frame temporal reasoning, and adaptive gated fusion to aggregate complementary observations for canonical geometry generation and consistent frame-wise object-state prediction. We further construct CoScene4D, a CARLA-based 4D dataset for structured semantic, geometric, spatial, and temporal supervision. CoMO4D is trained using CARLA-based CoScene4D together with nuScenes, while quantitative evaluation is conducted on the held-out V2X-Sep benchmark. The resulting performance demonstrates strong cross-domain transfer for temporally consistent multi-object 4D generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.