acceptodds
Under review as a conference paper at ICLR 2027

DSS-Net: Downstream Task-Driven Multi-modal Medical Image Fusion via Decoupled Spatial-Spectral Interaction

Abstract

Efficient multi-modal medical image fusion is vital for precision diagnostics. However, existing methods face a dilemma: models focusing solely on image fusion often prioritize visual aesthetics while neglecting semantic details required for downstream task. Meanwhile, unified fusion-downstream networks often suffer from feature entanglement and progressive semantic attenuation due to insufficient cross-task interaction. To address this, we propose DSS-Net, a novel framework that harmonizes high-fidelity fusion with downstream task performance. Central to our architecture are two specialized modules designed to reconcile representational conflicts: the Spatial Complementary Interaction (SCI) and the Global Spectral Integration (GSI). Specifically, SCI enables inter-modal complementary modulation to extract spatial visual features, while GSI executes spectral-domain aggregation to derive modality-invariant semantic representations. These decoupled features are then processed by dual-path decoders to achieve high-fidelity fusion and superior downstream utility. By decoupling representations into vision-driven spatial enhancement and semantic-driven spectral alignment, DSS-Net preserves intricate anatomical details while boosting downstream effectiveness. Extensive experiments demonstrate state-of-the-art performance in fusion quality and downstream task accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.