What Should 3D-LMMs See? Scene-State Curation for 3D Large Multimodal Models
Abstract
Recent progress in unified 3D Large Multimodal Models (3D-LMMs) has benefited substantially from increasing training-side investment, including large-scale 3D vision-language supervision, stronger cross-modal alignment, and multi-stage training. These advances enable models to acquire rich perceptual evidence from complex 3D scenes. Yet this evidence must still be organized into a shared scene state before it can support downstream reasoning. We identify a systematic mismatch in this process: scene-state construction uses candidate-level perceptual cues to decide which evidence should be prioritized or treated as redundant, but these cues do not necessarily align with the physical structure that the scene state should preserve. In particular, saliency is systematically coupled with physical support, so large-support candidates tend to receive greater representation priority, while semantic similarity alone can cause spatially distinct instances to be treated as redundant, suppressing independent evidence. We therefore propose **Physics-Aware 3D Scene Representation Curation (PARC-3D)**, a training-free framework for scene-state curation. PARC-3D calibrates the support-associated saliency advantage and models redundancy jointly through semantic and spatial criteria, yielding scene states that better preserve relevance, complementarity, and physical separability. Evaluations on referring segmentation, 3D question answering, and dense captioning show gains across all three task families while holding model parameters and scene-state size fixed, with the largest improvements on referring segmentation. These results highlight scene-state curation as a distinct representation axis complementary to data scaling, representation learning, and cross-modal optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.