acceptodds
Under review as a conference paper at ICLR 2027

Medical Image Alignment Assessment as a Test of Generalist Visual Reasoning in Frontier Multimodal Models

Abstract

Frontier multimodal large language models (MLLMs) are increasingly positioned as general purpose visual reasoners as part of the quest for artificial general intelligence. A key test of this generality is whether they can perform novel visual judgments that humans can make reliably from visual evidence and task instructions, without task-specific parameter optimisation. We investigate this question through the task of medical image alignment assessment, where the goal is to establish whether there is anatomical correspondence between two images. Assessing spatial alignment is a fundamental component of medical imaging pipelines across within-subject, between-subject, and cross-scanner settings. Human visual assessment of image alignment is still the gold standard and most common approach; however, it requires trained operators and is impractical to scale for large datasets. Existing automated methods using conventional image quality metrics or convolutional neural networks (CNNs) are generally designed for particular modalities, anatomical regions, or alignment tasks, not generalisable readily to new settings. We evaluate recent generations of MLLMs on two exemplar medical image alignment tasks, varying both prompting strategies and image-presentation methods. We compare against a smaller, locally fine-tuned MLLM and a task-specific CNN, allowing us to examine the trade-off between frontier general purpose models that cannot be used locally and smaller models that require specific task optimisation but can be used locally. Our results show that we are reaching an inflection point, where frontier MLLMs can now effectively perform visual assessment of medical image alignment. Models released only a few months ago generalise poorly and, in some settings, perform barely above chance, whereas GPT-6 achieves over 85% across almost all scenarios tested, including under zero-shot evaluation. Fine-tuned local models can match or exceed frontier-model performance on the tasks on which they are trained, but transfer substantially less effectively to unseen settings. These findings identify medical image alignment as a useful test bed for generalist visual reasoning. They also suggest that frontier multimodal models are beginning to acquire transferable visual assessment capabilities that could support a common quality-control mechanism across heterogeneous medical-imaging pipelines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.