acceptodds
Under review as a conference paper at ICLR 2027

Structure as State, Not Skip: Bridging Semantic Representation and Structural Detail for 3D Medical Image Segmentation

Abstract

Transformer-based architectures have become widespread in 3D medical image segmentation, yet their gains over strong convolutional baselines remain limited. We trace this mismatch to a representation–structure dilemma: skip connections preserve structure but let information bypass the Transformer; removing them makes the Transformer indispensable but weakens structural detail. The dilemma stems from conflating a task requirement with a U-Net design choice: structure must be preserved, but it need not pass through skip connections. Decoupling the two reveals a third possibility: structure as state, not skip. We realize this principle with Flumina, a multi-stream network that maintains multi-scale persistent states within a shared Vision Transformer backbone. Each sublayer combines these states, applies a shared transformation, and produces token-level updates for each state, while identity residual mappings maintain them across depth. Flumina thus keeps the Transformer responsible for representation learning while preserving structural detail. Experiments across multiple benchmarks and random seeds demonstrate its effectiveness. Code is available at the anonymous link.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.