SAFE-Diff: Where Anatomy Meets Frequency in Virtual Contrast Enhancement
Abstract
SAFE-Diff: Segmentation-Conditioned and Frequency-Aware Diffusion for Anatomy-Preserving Contrast Synthesis Download PDF Yuankai Wang, Jianyang Xie, Feixiang Zhou, Zhongli Wu, Tim Fairbairn, Gregory Yoke Hong Lip, John K Field, Yalin Zheng 20 Jul 2026 (modified: 20 Aug 2026) AAAI 2027 Conference Submission Conference, Area Chairs, Senior Program Committee, Program Committee, Authors Revisions CC BY 4.0 TL;DR: SAFE-Diff is a segmentation-conditioned, frequency-aware diffusion framework for synthesizing contrast-enhanced medical images from non-contrast scans while preserving fine anatomical structures. Abstract: Angiography is a contrast-enhanced imaging technique that improves the visualization of vascular structures, thereby facilitating the detection and assessment of vascular dis- eases. but many angiographic examinations rely on con- trast agents that may pose risks to susceptible patients. Vir- tual angiography synthesis from non-contrast scans may re- duce this reliance; however, existing medical image trans- lation methods can still have difficulty preserving fine vas- cular structures, occasionally producing blurred, displaced, or spurious details despite visually plausible contrast pat- terns.To address these challenges, we propose SAFE-Diff, a segmentation-conditioned and frequency-aware diffusion framework that combines global cross-modal alignment with local anatomical guidance. First, modality-specific VAE en- coders and a contrastive projection objective align paired source and target representations at the case level, provid- ing a global cross-modal condition for generation. Second, source-derived segmentation masks and outline maps are in- corporated into each denoising step to encourage consis- tency with the source anatomy. A multi-band spectral at- tention module further modulates bottleneck features using frequency-domain statistics to improve the representation of fine vessels and anatomical boundaries. We evaluate SAFE- Diff on 1,600 paired non-contrast CT and CTA volumes and on brain T1-to-T1ce synthesis, comparing it with represen- tative paired and unpaired translation methods. SAFE-Diff achieves 25.1 dB PSNR, 0.77 SSIM, and 0.26 LPIPS on CT- to-CTA synthesis, and 29.7 dB PSNR, 0.87 SSIM, and 0.19 LPIPS on T1-to-T1ce synthesis, obtaining the best or tied- best performance across the reported metrics. Component ab- lations, downstream structural evaluation, and robustness ex- periments further support the effectiveness of cross-modal alignment, structural conditioning, and spectral modulation. These findings provide evidence that SAFE-Diff improves global fidelity and proxy structural consistency on the evalu- ated paired datasets, although external and clinical validation remains necessary.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.