AnyGranGS: Agentic 3DGS Segmentation with Versatile Granularity
Abstract
3D Gaussian Splatting (3DGS) enables language-guided 3D scene understanding, while practical queries may target varying semantic granularities. Recent agentic methods leverage MLLMs for target reasoning and refinement, yet versatile-granularity segmentation remains challenged by (i) inflexible view acquisition, where predefined views may insufficiently observe fine-grained targets, and (ii) cross-view spatial ambiguity, where object-centric spatial relations can be inconsistently interpreted across viewpoints. We propose **AnyGranGS**, an agentic framework addressing both challenges. First, **Target-centric Adaptive Observation (TAO)** locates a coarse target proxy and constructs target-specific virtual views, while **Visibility Gain-based Selection (VGS)** selects complementary candidate views based on Gaussian visibility gain. Second, **Semantic Axis Constraints (SAC)** establish object-centric semantic axes for consistent part localization across views. Multi-view masks are then aggregated through visibility-weighted Gaussian evidence for target refinement. We further construct **Gran-LERF** and **Gran-3D-OVS** for versatile-granularity evaluation. Experiments on Gran-LERF, Gran-3D-OVS, and Ref-LERF demonstrate the effectiveness of AnyGranGS across diverse query granularities, improving the average mIoU over the strongest baselines by **15.37, 21.61, and 12.58 points**, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.