acceptodds
Under review as a conference paper at ICLR 2027

ReTax-Seg: Visually Grounded Reversible Taxonomy Transport for Image Segmentation

Abstract

Existing taxonomy-aware segmenters reconcile incompatible label names but leave a fundamental ambiguity: a dataset label can be a policy-dependent union of visual concepts, yet its mapping to a unified vocabulary remains global and one-way. A fixed mapping cannot identify the constituent present in a region. A language-derived mapping imports name semantics without observing that region. A learned latent vocabulary can also collapse despite correct final-label predictions. ReTax-Seg addresses these failures jointly. A visual decoder predicts 39 fine-grained semantic atoms, and an end-to-end transport field compiles them into native labels with a different atom–class distribution at every spatial location. Class names supply only a learned-strength CLIP prior, while supervised visual features and reliability-gated MobileSAM geometry determine the local correction. Reverse transport requires predicted native labels to reconstruct the visual atom distribution, making semantic preservation an explicit training objective. To our knowledge, ReTax-Seg first combines pixel-conditioned class–atom transport, visual-over-language arbitration, and bidirectional taxonomy consistency in one differentiable compiler. Six independent checkpoints average 73.63% test mIoU. One jointly trained six-dataset checkpoint, ReTax-Seg-C, shares the atom representation and reaches 74.05%, outperforming the compared methods on every dataset.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.