acceptodds
Under review as a conference paper at ICLR 2027

Where to Refine and What to Recover: Disentangled Cross-modal Semantic Alignment for CT-to-PET Generation

Abstract

Positron Emission Tomography (PET) and Computed Tomography (CT) play complementary and essential roles in advanced healthcare, particularly for cancer diagnosis. However, CT is far more widely available than PET, primarily due to its substantially lower acquisition and operational costs. To bridge this gap, many CT-to-PET generation methods are developed based on structural alignment while overlooking the intrinsic characteristics of PET, such as metabolic activity, which limits the clinical utility of the synthesized PET images. To address this challenge, we propose a disentangled alignment framework that formulates CT-to-PET generation as a controlled residual generation problem, explicitly guiding both the spatial location ("where") and semantic content ("what") of PET signals. Specifically, a frozen backbone first generates a stable PET prediction, while a Scalable Interpolant Transformer (SiT) models the remaining residual refinement. Then, to characterize what should be recovered, intermediate SiT representations are adopted to align the semantic residual via deep supervision, thereby reinforcing the learning of missing semantics. Furthermore, an activation gate is introduced to identify regions requiring refinement while suppressing redundant residual noise in well-reconstructed areas. Extensive experiments on the popular Lung-PET-CT-Dx and RIDER-Lung-PET-CT datasets demonstrate that our method significantly outperforms state-of-the-art baselines, highlighting its strong potential for CT-to-PET image generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.