Integration by Parts: Compositional Multimodal Learning for Dynamic 3D Scene Graph Generation
Abstract
Three-dimensional (3D) scene understanding is essential for autonomous systems operating in dynamic environments. While RGB, LiDAR, and textual semantics provide complementary information, existing dynamic 3D scene graph methods often rely on holistic object representations that are sensitive to occlusion, viewpoint changes, sensor noise, incomplete observations, and cross-modal misalignment. We propose Integration by Parts (IBP), a compositional multimodal framework that represents objects using fine-grained latent components rather than single embeddings. Shared learnable queries decompose RGB and LiDAR observations into latent components, which are explicitly aligned across modalities and integrated through recomposition with global multimodal and textual context to form compositional object representations. A temporal module associates these representations across observations and maintains semantic consistency for spatial-temporal reasoning. Experiments on KITTI-360 demonstrate the effectiveness of IBP for multimodal alignment, temporal reasoning, and robustness to incomplete observations. The code is available at: https://anonymous.4open.science/r/ibp_anonymized-/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.