acceptodds
Under review as a conference paper at ICLR 2027

GeMI3D: Geometry-Grounded Modality-Independent Indoor 3D Object Detection

Abstract

Indoor 3D object detection localizes and recognizes objects in 3D space from images, depth maps, point clouds, or their combinations, but existing detectors are typically designed for predefined input configurations. When the available inputs change, these detectors may require architectural changes or additional training. We therefore study modality-independent indoor 3D object detection, where a single detector handles any combination of image, depth, and point-cloud inputs. To this end, we propose GEMI3D, a geometry-grounded modality-independent framework for indoor 3D object detection, which maps heterogeneous inputs into a stable 3D geometric hub. First, Cross-Modal Hub Construction learns a stable spatial representation from point-cloud data and then extends the geometry-grounded hub with depth and image inputs mapped into the same 3D space, while feature anchoring preserves the point-cloud representation. Second, Modality-Conditioned Hub Adaptation uses modality dropout to expose the detector to different input combinations and fuses the available hub features for a modality-agnostic detection head. To comprehensively evaluate modality independence with inputs from distinct sources, we further construct TriScan3D, a frame-level benchmark derived from ScanNet++. With all modalities, GEMI3D achieves 68.6/52.3 mAP25/mAP50 on SUN RGB-D and 70.7/51.4 on TriScan3D. Across all seven non-empty input combinations, it maintains robust detection performance without retraining. The code will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.