MulGeCa: Anchoring Scattered Semantic Representations to Physical Entities
Abstract
Recent open-vocabulary segmentation methods leverage pretrained vision-language models (VLMs) to recognize arbitrary text queries beyond the constraints of closed-set categories. Despite their promise, adapting image-level vision-language representations to dense pixel-level prediction remains fundamentally challenging. Semantic representations from VLMs are often scattered and weakly structured, making it difficult to bind open-vocabulary concepts to their corresponding physical entities. To solve this problem, our key insight is that category-agnostic depth cues provide spatial constraints that can bind semantic responses to depth-consistent physical support. To this end, we propose MulGeCa, which leverages estimated relative depth to hierarchically calibrate scattered VLM semantics at the scene, region, and instance levels. Specifically, MulGeCa incorporates a global-geometry text embedding calibrator to inject dynamic geometric context into text prompts, a geometric-contextual anchor aggregator to modulate attention mechanisms with geometric-consistency biases, and a centroid-purified mask refiner to suppress boundary noise via centripetal weighted aggregation. Extensive experiments on challenging benchmarks demonstrate that MulGeCa achieves superior performance, with particular improvements in cluttered scenarios. Notably, our method exhibits strong robustness in complex scenes, providing an effective pathway for anchoring scattered semantic representations to depth-consistent regions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.