acceptodds
Under review as a conference paper at ICLR 2027

MonoVoc: Efficient Object-Centric Open-Vocabulary 3D Gaussian Mapping from Monocular Video

Abstract

Open-vocabulary 3D scene understanding lets users query reconstructed environments in natural language, but existing Gaussian-based methods either store dense per-Gaussian language features or couple semantics to an optimization-heavy reconstruction, both of which scale poorly on ordinary monocular video. We present MonoVoc, an object-centric framework that turns a single monocular walkthrough and its segmentation masks into a compact semantic Gaussian map paired with an object-level language database. MonoVoc attaches three learnable semantic coefficients to each Gaussian of a fixed SLAM backbone, recovers per-Gaussian semantic evidence by windowed deblending restricted to projected Gaussian footprints, and resolves discrete object identities through confidence-gated quantization that combines palette agreement with spatial consistency. Each object is then described once from a set of diverse keyframes and encoded into a single shared embedding, adding semantics on top of reconstruction without a separate language-feature distillation stage and reducing language storage from O(ND) to O(N + KD) for N Gaussians and K objects. On Replica and ScanNet, MonoVoc produces a queryable semantic map in 13–13.5 minutes end-to-end, against 186 minutes for ObjectGS, using ≈160K Gaussians and 14 MB; 84% and 91% less memory than ObjectGS and SceneSplat. Built on HI-SLAM2, MonoVoc inherits its rendering quality and reaches, 91.65 mIoU on Replica given ground-truth 2D segmentation, 84%/93% Top-1/Top-3 object retrieval, and 86% accuracy on a controlled scene-grounded QA benchmark, while ObjectGS retains an edge on object removal. A 34-participant study prefers our maps in 52.9% of responses.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.