CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis
Abstract
Generating an image for a requested medical condition requires more than plausible appearance: the condition must be visible in the image. Medical-domain adaptation, richer descriptions, and reward or preference alignment improve different aspects of generation, but agreement with text or visual plausibility alone does not necessarily make the requested condition recognizable in a new sample. We ask whether direct feedback on generated images improves expression of the requested condition beyond equal additional text-guided reconstruction. CRAFT (Clinical Reward-Aligned Finetuning for Medical Image Synthesis) tests this question by retaining reconstruction and scoring sampled images against an instance description, class-shared visual criteria, and the target label using a classifier trained on real images. In a matched CheXpert continuation from the same TI+LoRA initialization, CRAFT exceeds reconstruction-only training in DINOv2 balanced accuracy by 10.70 percentage points after five additional epochs [paired source-bootstrap 95% CI 7.28, 14.12]. Across three CheXpert training runs, removing diagnostic feedback lowers balanced accuracy for label recognition, while MetaCLIP2 description and checklist similarities barely change. Across four domains, CRAFT has the highest synthetic-method point estimates for DINOv2 balanced accuracy and both MetaCLIP2 similarities. Two physicians selected CRAFT in 55.3% versus TI+LoRA in 26.3% of 800 method-anonymous four-way judgments. Feedback on sampled images thus makes requested conditions more recognizable to a held-out classifier trained on real images while improving alignment with instance descriptions and shared clinical criteria.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.