acceptodds
Under review as a conference paper at ICLR 2027

Sub-Semantic Image Segmentation

Abstract

Semantic segmentation tells us what a region is, but many useful visual distinctions depend on how it looks. Bare and moss-covered rock, peeling paint and exposed wood, or grass growing through sand may share semantic labels while forming clearly distinct appearance regions. Traditionally, segmentation partitions scenes into predefined objects and an amorphous “stuff” category. However, as modern segmentation models increasingly integrate with natural language, this rigid di- chotomy becomes obsolete, and there is no longer a reason to discard rich visual information into a generic “stuff” bucket. Yet, driven by a strong semantic bias, these advanced models degrade when encountering regions that are too abstract for a fixed dictionary noun, but visually distinct enough to elicit rich human descrip- tions. To provide a more fitting representation and explicitly target this vulnerability, we introduce sub-semantic image segmentation, the task of autonomously discov- ering such regions, describing each in free-form language and assigning every description an associated mask, all without receiving region prompts or the number of regions in advance. To make this problem measurable, we contribute three com- plementary task-specific benchmarks spanning annotation-derived natural scenes with variable region counts, independently constructed natural surface transitions and entirely new controlled synthetic compositions, together with validation on an established real-world texture dataset. We propose Detecture, an integrated description-to-partition system that generates structured appearance descriptions, extracts isolated grounding states, predicts description-conditioned masks and resolves them through pixelwise competition. Across all four evaluation routes, Detecture consistently outperforms native GLaMM and current Sa2VA in region covering, and leads the evaluated native language-and-mask model class on both natural benchmarks. Controlled interventions also show that descriptions actively guide region assignment: changing appearance wording alters masks, and mask identities follow appearance edits substantially more often than spatial edits among jointly eligible cases. Together, these results establish language as a practical interface for discovering and addressing visual structure beyond fixed category vocabularies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.