acceptodds
Under review as a conference paper at ICLR 2027

MANY VIEWS, FEW INSIGHTS: RETHINKING THE ALIGNMENT OF VISION LANGUAGE MODEL FOR CLINICAL MULTI-IMAGES ANALYSIS

Abstract

Clinical diagnosis routinely requires synthesizing evidence across multiple images of the same patient, yet current vision-language models (VLMs) fall far short, largely because they fail to integrate evidence across images. The standard alignment recipe to align VLM on clinical multi-images is supervised fine-tuning (SFT) optionally followed by reinforcement learning (RL). This baseline recipe carries several under-studied free choices, including whether to add an RL stage, the SFT epoch count, and the LoRA rank. We systematically compare these choices, evaluating each configuration on in-domain and out-of-domain clinical benchmarks, and distill several insightful findings showing that existing recipes leave substantial performance on the table for assessing multiple images. We trace this gap to a specific training-time failure: SFT-only models show little hallucination on in-domain benchmarks but hallucinate substantially more on out-of-domain benchmarks, pointing to non-visual shortcuts that break under distribution shift. We call this supervision-modality mismatch. Building on these findings, we propose MULTI-MULTI for clinical multi-modal and multi-image analysis, which decomposes multi-image reasoning into an explicit, per-image-grounded structure and introduces Image-Swap Supervision (ISS) which pairs a case's images with a mismatched clinical history and mixed into fine-tuning samples so each reasoning step is recoverable only from the images. Despite this lightweight recipe, the resulting fine-tuned 8B open-source model reaches in-domain accuracy comparable to zero-shot open models two orders of magnitude larger, and closes part of the gap to leading proprietary models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.