AnyCT-Ω: Scaling Cross-View Representations for Heterogeneous Sparse-View CT Reconstruction
Abstract
Sparse-view computed tomography (CT) reconstruction aims to recover a volumetric attenuation field from a small number of X-ray projections. Existing learning-based methods are often optimized for a single dataset or anatomical domain, which limits their robustness under changes in anatomy, preprocessing, and acquisition statistics. We present AnyCT-, a unified framework for heterogeneous sparse-view CT reconstruction that adapts transferable cross-view representations from a pretrained VGGT- backbone to projection-based medical imaging. The model uses known CT geometry to align multi-view features with arbitrary three-dimensional query points and employs a content-routed residual mixture-of-experts head to adapt dense features across heterogeneous inputs. We further construct AnyCT-70K, a standardized corpus containing 73,610 processed samples from seven public CT sources covering thoracic, abdominal, pelvic, and musculoskeletal anatomy. Under a common joint-training protocol, AnyCT- outperforms the strongest baseline on the mixed AnyCT-70K benchmark by 1.50, 1.72, and 1.64 dB PSNR at 6, 8, and 10 views, respectively. The corresponding SSIM improvements are 3.76, 3.63, and 3.06 points. Additional experiments show positive transfer when increasing the number of training sources and demonstrate that the proposed model reaches competitive reconstruction quality with 30,000 optimization steps. These results suggest that geometry-guided adaptation of transferable cross-view representations is a promising direction for heterogeneous sparse-view CT reconstruction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.