acceptodds
Under review as a conference paper at ICLR 2027

Ophth-500K: Curating Unstructured Video Streams for Scaling Multimodal Ophthalmology Foundation Models

Abstract

Large-scale vision-language learning in ophthalmology remains constrained by the limited availability of multimodal data with clinically descriptive supervision. We introduce Ophth-500K, a large-scale multimodal ophthalmology dataset constructed from 14,769 hours of expert-led videos. Our automated data engine identifies clinically relevant ophthalmic content, aligns visual segments with expert narration, and converts continuous video streams into structured image, text, and audio samples. Ophth-500K contains 467,697 aligned multimodal samples spanning Optical Coherence Tomography, Color Fundus Photography, and Ultra Widefield imaging. Its source content covers 116 languages and more than 100 countries and regions, while UMLS-based analysis identifies 1,885 anatomical concepts, 689 pathologic findings, and 3,379 disease concepts. To evaluate the learning value of the curated supervision, we train OphthCLIP using a standard contrastive vision-language objective. OphthCLIP substantially improves cross-modal retrieval over general-purpose, medical, and ophthalmology-specific baselines, achieving 32.6% Recall@1 for image-to-text retrieval and 33.2% for text-to-image retrieval. These results demonstrate the value of expert ophthalmology videos as a scalable source of structured multimodal supervision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.