acceptodds
Under review as a conference paper at ICLR 2027

Integration by Parts: Compositional Multimodal Learning for Dynamic 3D Scene Graph Generation

Abstract

Three-dimensional (3D) scene understanding is essential for autonomous systems operating in dynamic environments. While RGB, LiDAR, and textual semantics provide complementary information, existing dynamic 3D scene graph methods often rely on holistic object representations that are sensitive to occlusion, viewpoint changes, sensor noise, incomplete observations, and cross-modal misalignment. We propose Integration by Parts (IBP), a compositional multimodal framework that represents objects using fine-grained latent components rather than single embeddings. Shared learnable queries decompose RGB and LiDAR observations into latent components, which are explicitly aligned across modalities and integrated through recomposition with global multimodal and textual context to form compositional object representations. A temporal module associates these representations across observations and maintains semantic consistency for spatial-temporal reasoning. Experiments on KITTI-360 demonstrate the effectiveness of IBP for multimodal alignment, temporal reasoning, and robustness to incomplete observations. The code is available at: https://anonymous.4open.science/r/ibp_anonymized-/

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.