RetinaRobust: Benchmarking Artifact Robustness and Adaptation in Retinal Foundation Models
Abstract
Retinal foundation models support diverse diagnostic tasks through the analysis of routine, non-invasive images of the eye. Nevertheless, these images are subject to a range of imaging artifacts and variations across clinical settings. We benchmark six retinal foundation models on 17 diagnostic tasks to understand their robustness to artifacts and color removal. All models exhibit performance degradation with the loss of image quality and chromatic information, and substantial gaps remain even after end-to-end fine-tuning for downstream tasks. We examine whether downstream adaptations, including linear probing and end-to-end fine-tuning, can be still utilized to recover the lost performance. We find that training with clean and artifact-corrupted images improves performance not only on artifact-corrupted test images but also on clean test images under both linear probing and end-to-end fine-tuning. These improvements, however, transfer only partly to unseen artifacts. Similarly, transfer of robustness gained from artifacts to the color domain, and vice versa, is limited. Therefore, retinal foundation models must be rigorously evaluated to ensure that their performance generalizes to scenarios with possible corruptions observed beyond training. Our results also motivate more principled machine learning solutions to address the broad-spectrum resilience of these models that facilitate diverse ophthalmic research and clinical applications.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.