acceptodds
Under review as a conference paper at ICLR 2027

Composing Complementary Dense Representation for Cost Aggregation in Open-Vocabulary Semantic Segmentation

Abstract

Open-vocabulary semantic segmentation enables dense scene understanding beyond a fixed training vocabulary. Existing cost-aggregation methods, however, typically reuse a single dense visual feature for both cost map construction and spatial guidance. Coupling these objectives within a shared dense feature space induces competition between semantic discrimination and spatial coherence, and simultaneously constrains the diversity of visual representations available across open-vocabulary scenarios. To address these limitations, we introduce Complementary Dense Representations (CoDR), a novel framework for constructing and composing dense representations for cost aggregation. Specifically, we first introduce Cost Spatial Decoupling, which assigns cost map construction and spatial guidance to distinct representation spaces, freeing both from the constraints of a shared feature. Building on this decoupling, Representation Diversification constructs multiple complementary representations within each space to preserve distinct semantic and spatial cues. Furthermore, we design Scene-Adaptive Representation Composition to hierarchically compose structurally distinct representations for each input scene, replacing a fixed global preference with scene-specific composition. Extensive experiments on five standard benchmarks demonstrate the strong performance of CoDR, which achieves 19.8 mIoU on ADE847. We will release all code and models upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.