Ophth-500K: Curating Unstructured Video Streams for Scaling Multimodal Ophthalmology Foundation Models
Abstract
Large-scale vision-language learning in ophthalmology remains constrained by the limited availability of multimodal data with clinically descriptive supervision. We introduce Ophth-500K, a large-scale multimodal ophthalmology dataset constructed from 14,769 hours of expert-led videos. Our automated data engine identifies clinically relevant ophthalmic content, aligns visual segments with expert narration, and converts continuous video streams into structured image, text, and audio samples. Ophth-500K contains 467,697 aligned multimodal samples spanning Optical Coherence Tomography, Color Fundus Photography, and Ultra Widefield imaging. Its source content covers 116 languages and more than 100 countries and regions, while UMLS-based analysis identifies 1,885 anatomical concepts, 689 pathologic findings, and 3,379 disease concepts. To evaluate the learning value of the curated supervision, we train OphthCLIP using a standard contrastive vision-language objective. OphthCLIP substantially improves cross-modal retrieval over general-purpose, medical, and ophthalmology-specific baselines, achieving 32.6% Recall@1 for image-to-text retrieval and 33.2% for text-to-image retrieval. These results demonstrate the value of expert ophthalmology videos as a scalable source of structured multimodal supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.