OphMem: Leveraging Cross-Case Experience to Mitigate Visual Ambiguity in Ophthalmic Surgical Video Understanding
Abstract
Ophthalmic surgical video understanding can assist clinicians in analyzing surgical workflows, identifying abnormalities, and supporting postoperative assessment. However, a key challenge lies in visual ambiguity, where visually similar observations may correspond to different surgical states. Existing methods mainly rely on intra-case contextual information, which does not always provide reliable discriminative cues. To address this limitation, we propose OphMem, a framework that continually accumulates and reuses cross-case procedural experience. OphMem consists of two core modules: Experience-Conditioned Agentic Reasoning, which retrieves applicable experience to guide video inspection and tool use, and Persistent Experience Memory, which transforms completed reasoning trajectories into medically grounded reusable experience and consolidates them into an evolving experience bank. To further assess visual ambiguity, we construct a dedicated benchmark with three tasks: History-Conditioned Step Disambiguation, Ambiguous Complication Verification, and Ophthalmic Temporal VQA, together with a sequential experience protocol. Across the three tasks, OphMem achieves a mean relative gain of 8% in primary accuracy over various benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.