acceptodds
Under review as a conference paper at ICLR 2027

OphthoFlow: An Agentic Workflow for Clinical Decision Support in Ophthalmology

Abstract

Ophthalmic imaging guides the management of the leading causes of irreversible blindness and, because the retina offers a direct, non-invasive view of the microvasculature and central nervous system, also serves as an inexpensive marker of systemic health. Yet neither of the two dominant AI paradigms fits clinical workflows well. Vision foundation models produce predictions they cannot reason about, and large language models (LLMs) reason fluently but hallucinate and cannot analyze images reliably. We present OphthoFlow, an agentic workflow for clinical decision support in ophthalmology. An LLM orchestrator, OphthoAgent, plans and calls OphthoTools, a suite of 61 validated task-specific models covering six clinical functions across four ophthalmic imaging modalities, and grounds its recommendations with OphthoRAG, a hybrid retrieval module over a curated ophthalmology knowledge base, behind a clinician-facing interface, OphthoChat, that exposes every step as an auditable trace. The orchestrator is text-based, so any LLM can serve as its backbone and raw images never leave the local deployment. To evaluate the workflow, we build OphthoBench, a benchmark spanning 57 datasets that separately measures clinical tool accuracy, end-to-end clinical visual question answering (VQA) over 25,000 items, and clinical workflow orchestration over 14,000 queries from 56 clinical tasks with 24 evaluation metrics, complemented by a robustness evaluation on tool-removal scenarios and on an adversarial set of out-of-scope images and ungradable retinal images where the correct response is to abstain or to warn, and a fairness analysis across demographic groups. With a GPT-5.3 backbone, OphthoAgent answers within 1.1 percentage points of a tool-oracle model that always calls the correct tool, exceeds the best of nine end-to-end vision-language models (VLMs) by 8.4 macro-F1 points, identifies the required tool set with an F1 of 92.3, and abstains with a correct reason on 99% of out-of-scope images.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.