acceptodds
Under review as a conference paper at ICLR 2027

MedVision-X: Unifying Medical Image Analysis and Synthesis through Generative Modeling

Abstract

Medical image analysis and synthesis are served by separate specialist models: segmentation networks predict label maps, restoration networks recover intensities, and translation or editing networks synthesize new images, each with its own architecture, objective, and output head. We ask how far a single generative model with one image output pathway can go across these heterogeneous tasks. We present , which casts segmentation, super-resolution, cross-modal translation, virtual immunohistochemical staining, CT denoising, and lesion editing as instruction-conditioned image generation. Masks are rendered as color-coded images with an optional text legend and continuous targets are generated directly, so every task is trained with the same flow-matching objective and decoded by the same image decoder, without task-specific heads or losses. Starting from the SenseNova-Vision unified multimodal model, is adapted in two stages: broad adaptation on 1.04M medical records, followed by balanced joint training on the training folds of 16 benchmarks. Across 14 test benchmarks and 62 scored baseline entries, obtains the best primary endpoint on 8 and ranks second on 2 more. The gains are largest for physically meaningful targets: one checkpoint reduces body HU MAE from 99.0 to 39.1 on low-dose CT and from 83.1 to 55.2 on MRI-to-CT synthesis, and improves PSNR by 4.0 dB on LoDoPaB-CT and 3.2 dB on IXI. Dataset-specific segmentation checkpoints lead on BUSI and Kvasir-SEG, while ISIC, DRIVE, and pixel-fidelity super-resolution remain behind the strongest specialists. A mask round-trip diagnostic shows that the image representation preserves thin structures, locating the remaining gap in conditional spatial precision rather than in the unified output format.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.