VOLT: 3D Medical Foundation Model for Volumetric Latent Tokenization
Abstract
Medical foundation model (FM) development faces a critical trade-off between semantically aligned multimodal models constrained by paired data and scalable scan-only models with weaker clinical grounding, compounded by volumetric anisotropy and the scarcity of paired scan–text data. To bridge this gap, we introduce VOLT, a scan-only 3D medical FM for high-fidelity volumetric tokenization trained through a two-stage strategy. Stage-1 learns robust through-plane priors through Depth Shuffle reconstruction and discretizes latents into compact codewords using the proposed Additive Binary Spherical Quantization (ABSQ) module. Stage-2 aligns these representations with clinical semantics through patch-consistency matching and distillation from a frozen vision-language teacher, without requiring paired reports for the training scans. VOLT yields semantically structured representations that transfer effectively across segmentation, classification, visual question answering (VQA), report generation, and synthetic data generation. Compared with the strongest reported baseline for each metric, VOLT achieves relative improvements of 13.57% in PSNR and 6.22% in SSIM for volumetric reconstruction, 0.35% in Dice and a 10.98% reduction in HD95 for segmentation, and 1.84% in mean VQA accuracy. For report generation, it improves BLEU, METEOR, and BERTScore by 4.47%, 2.11%, and 0.52%, respectively. Code and dataset will be open-sourced.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.