Spatially Structured Codebook Learning for 3D Instance-Level Perception of Aerial-Ground Scenes
Abstract
Recent advances in unified aerial–ground reconstruction have enabled photorealistic modeling of large-scale environments, yet fine-grained scene understanding remains underexplored. Achieving consistent instance-level semantics is particularly challenging due to sparse per-view observations over a large identity space and severe cross-view variations in projected scale and visible detail. These factors complicate scene-wide optimization and lead to over- or under-segmented 2D instance masks with inconsistent supervision. To address these challenges, we propose Spatially Structured Codebook Learning (SSCL), a scalable framework for instance-level aerial–ground scene understanding. SSCL first organizes the scene into spatial root regions, each with a compact codebook of leaves representing local instance identities, and jointly learns cross-root connectivity to fuse leaves into globally consistent instances. A granularity-aware codebook learning method is further introduced to heuristically mitigate conflicting instance supervision. We further introduce Aerial–Ground Segmentation (AG-Seg), a new benchmark dataset for evaluating instance segmentation and open-vocabulary retrieval in aerial–ground scenes. Extensive experiments on AG-Seg demonstrate substantial improvements over state-of-the-art methods in both tasks, enabling fine-grained interaction with large aerial–ground environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.