Can Multimodal Large Language Models Understand OCT?
Abstract
Optical coherence tomography (OCT) is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have shown strong potential in medical image analysis, existing benchmarks mainly focus on coarse-grained disease classification or visual question answering, failing to comprehensively evaluate the cognitive process from visual perception to clinical reasoning. To address this gap, we present OCT-Bench, a comprehensive benchmark for OCT image understanding. OCT-Bench contains 10,076 high-quality multiple-choice questions derived from 4,137 OCT images across seven public datasets. Following the clinical interpretation workflow, we establish a hierarchical taxonomy of 20 fine-grained tasks spanning three capability dimensions: Perception, Cognition, and Reasoning. These tasks cover imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, treatment decision-making, and prognosis. We evaluate 20 representative MLLMs, including close-source, open-source, and medical-domain models. Results show that current models remain far from reliable OCT understanding, while neither medical-domain adaptation nor larger model scale consistently improves performance across capability levels. OCT-Bench provides a comprehensive benchmark for evaluating OCT understanding and facilitates the development of clinically grounded MLLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.