CameraTime Factorization for Unified Multi-Camera 4D Reconstruction
Abstract
Multi-camera 4D reconstruction is important for embodied perception. Synchronized robot observations naturally form a cameratime structure. Cameras at the same timestamp provide complementary geometric views, while observations from the same camera capture temporal evolution. Existing methods often flatten these observations into a single sequence or process each camera independently, and thus do not explicitly model these two relations. In this paper, we first introduce a CameraTime Formulation for embodied 4D reconstruction. According to formulation, we build RoboCT4D, a unified benchmark that standardizes cameratime observations and evaluation across simulated and real robot manipulation datasets. We then propose CxT4D with CameraTime Factorization. It encodes camera and temporal identities separately and alternates global attention across synchronized cameras and along individual camera streams. This design explicitly models cross-camera geometry and temporal correspondence. We further propose a Long-Sequence MV4D Strategy that reduces memory cost while preserving reconstruction accuracy. Experiments on four RoboCT4D subsets show that CxT4D improves depth and cross-camera rotation accuracy over existing methods. Our long-sequence inference reduces peak memory by 51.0% at eight timestamps and 60.3% at twelve timestamps, and enables longer inputs that otherwise run out of memory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.