acceptodds
Under review as a conference paper at ICLR 2027

Vision Transformer-Conditioned UNet for Domain-Adaptive Semantic Segmentation

Abstract

Despite significant advances in natural domain vision, biomedical semantic segmentation remains challenged by severe label scarcity, low signal-to-noise ratios, and fine anatomical structures. We introduce ViTC-UNet, a framework that bridges this gap by conditioning a UNet pixel decoder on frozen Vision Transformer (ViT) representations via a multi-stage, two-way attention decoder. Semantic identity is decoupled from output geometry using target-specific structure tokens, implemented either as learnable class vectors or embeddings from a language model. This design unifies the global representational capacity of large-scale ViTs with the spatial inductive biases and high-resolution decoding precision of UNets, bypassing computationally intensive and data-hungry Transformer fine-tuning. ViTC-UNet consistently outperforms baselines across CT and MRI modalities in both closed- and open-vocabulary settings, as well as out of distribution, demonstrating that structure-conditioned UNet decoding provides a sample-efficient, highly adaptable pathway for biomedical foundation models. Code will be released upon publication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.