Factorized Probabilistic Inference for Multimodal Few-Shot Point Cloud Segmentation
Abstract
Multimodal Few-Shot Point Cloud Semantic Segmentation (MFS-PCS) leverages Vision-Language Models (VLMs) to reduce reliance on dense 3D annotations. However, existing methods often perform inference in an entangled latent space that jointly mixes visual, textual, and geometric cues. Such a monolithic treatment obscures the distinct roles of semantic priors and geometric structure, often leading to biased prototypes and misalignment between external semantics and the intrinsic manifold of the query scene. From a probabilistic perspective, we argue that few-shot inference can be naturally factorized into semantic likelihood estimation and geometric manifold regularization. Motivated by this view, we propose Decoupling Semantic Likelihood and Geometric Priors (DSG). For semantic transfer, we reformulate support prototype construction as a Maximum Likelihood Estimation problem rather than a deterministic averaging procedure, and employ Expectation Maximization to estimate robust semantic prototypes from sparse and noisy support observations. For geometric regularization, we introduce unsupervised structural anchors that capture the latent Voronoi structure of the query manifold. These anchors act as transductive geometric priors, encouraging predictions to respect local neighborhood structure and remain coherent across geometrically related regions. By separating what a category is from where and how it appears in the query geometry, DSG better aligns external multimodal semantics with internal 3D topology. Extensive experiments on S3DIS and ScanNet show that DSG achieves new state-of-the-art performance for MFS-PCS.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.