Structure as State, Not Skip: Bridging Semantic Representation and Structural Detail for 3D Medical Image Segmentation
Abstract
Transformer-based architectures have become widespread in 3D medical image segmentation, yet their gains over strong convolutional baselines remain limited. We trace this mismatch to a representation–structure dilemma: skip connections preserve structure but let information bypass the Transformer; removing them makes the Transformer indispensable but weakens structural detail. The dilemma stems from conflating a task requirement with a U-Net design choice: structure must be preserved, but it need not pass through skip connections. Decoupling the two reveals a third possibility: structure as state, not skip. We realize this principle with Flumina, a multi-stream network that maintains multi-scale persistent states within a shared Vision Transformer backbone. Each sublayer combines these states, applies a shared transformation, and produces token-level updates for each state, while identity residual mappings maintain them across depth. Flumina thus keeps the Transformer responsible for representation learning while preserving structural detail. Experiments across multiple benchmarks and random seeds demonstrate its effectiveness. Code is available at the anonymous link.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.