MANY VIEWS, FEW INSIGHTS: RETHINKING THE ALIGNMENT OF VISION LANGUAGE MODEL FOR CLINICAL MULTI-IMAGES ANALYSIS
Abstract
Clinical diagnosis routinely requires synthesizing evidence across multiple images of the same patient, yet current vision-language models (VLMs) fall far short, largely because they fail to integrate evidence across images. The standard alignment recipe to align VLM on clinical multi-images is supervised fine-tuning (SFT) optionally followed by reinforcement learning (RL). This baseline recipe carries several under-studied free choices, including whether to add an RL stage, the SFT epoch count, and the LoRA rank. We systematically compare these choices, evaluating each configuration on in-domain and out-of-domain clinical benchmarks, and distill several insightful findings showing that existing recipes leave substantial performance on the table for assessing multiple images. We trace this gap to a specific training-time failure: SFT-only models show little hallucination on in-domain benchmarks but hallucinate substantially more on out-of-domain benchmarks, pointing to non-visual shortcuts that break under distribution shift. We call this supervision-modality mismatch. Building on these findings, we propose MULTI-MULTI for clinical multi-modal and multi-image analysis, which decomposes multi-image reasoning into an explicit, per-image-grounded structure and introduces Image-Swap Supervision (ISS) which pairs a case's images with a mismatched clinical history and mixed into fine-tuning samples so each reasoning step is recoverable only from the images. Despite this lightweight recipe, the resulting fine-tuned 8B open-source model reaches in-domain accuracy comparable to zero-shot open models two orders of magnitude larger, and closes part of the gap to leading proprietary models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.