UniCTPET: Unifying 3D CT and PET Generation through Adaptive Structural Exchange
Abstract
Generative modeling offers a promising way to alleviate the scarcity of paired CT/PET data, yet existing methods typically address different generation tasks with separate models, limiting knowledge reuse across tasks. Unifying these tasks requires accommodating heterogeneous structural constraints and cross-modal dependencies that differ in direction and evolve throughout generation. We propose a unified 3D latent flow-matching framework that supports single-modality, structurally conditioned, CT-to-PET, and joint CT/PET generation within one model, without task-specific retraining. At its core, Time Task Structure Exchange (TTSE) adaptively regulates cross-modal structural interactions using modality feature states and flow time, accommodating one-way conditioning in CT-to-PET synthesis and bidirectional coordination in joint generation without explicit task representations. Complementarily, Skeletal Lesion Guidance (SLG) provides spatial control by injecting multi-scale skeletal and lesion conditions through lightweight encoders, without replicating the backbone encoder. Experiments on three whole-body PET/CT datasets comprising 1,599 paired scans demonstrate improved generation fidelity and downstream lesion segmentation over competing generative methods. Averaged across datasets, our framework improves PET PSNR in joint generation by 3.10 dB over the strongest competing baseline, while TTSE improves CT-to-PET PSNR by 2.15 dB over standard cross-attention. These findings highlight adaptive structural exchange as an effective mechanism for unifying generation tasks with heterogeneous conditions and modality dependencies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.