Vision Transformer-Conditioned UNet for Domain-Adaptive Semantic Segmentation
Abstract
Despite significant advances in natural domain vision, biomedical semantic segmentation remains challenged by severe label scarcity, low signal-to-noise ratios, and fine anatomical structures. We introduce ViTC-UNet, a framework that bridges this gap by conditioning a UNet pixel decoder on frozen Vision Transformer (ViT) representations via a multi-stage, two-way attention decoder. Semantic identity is decoupled from output geometry using target-specific structure tokens, implemented either as learnable class vectors or embeddings from a language model. This design unifies the global representational capacity of large-scale ViTs with the spatial inductive biases and high-resolution decoding precision of UNets, bypassing computationally intensive and data-hungry Transformer fine-tuning. ViTC-UNet consistently outperforms baselines across CT and MRI modalities in both closed- and open-vocabulary settings, as well as out of distribution, demonstrating that structure-conditioned UNet decoding provides a sample-efficient, highly adaptable pathway for biomedical foundation models. Code will be released upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.