acceptodds
Under review as a conference paper at ICLR 2027

Learning from Heterogeneous Annotation Mechanisms for Unified Retinal Prediction

Abstract

Learning a unified retinal predictor from datasets collected for different clinical purposes is complicated by a fundamental ambiguity: an unrecorded label is not necessarily clinically negative. Because annotations follow heterogeneous, source-specific protocols, flattening such datasets into a common binary label matrix can introduce systematic false-negative supervision. Masking unobserved entries prevents this error, but uses label observability only to gate the loss and leaves both the structured observation process and uneven supervision across source–label blocks unmodeled. We introduce Annotation-Mechanism-Aware learning (AMA), which separates clinical prediction from label observability. AMA combines an auxiliary label-observability objective with source–label schema balancing, while requiring neither source identity nor observation masks at inference. We evaluate AMA on 252,116 fundus photographs from 54 public datasets whose annotations are harmonized into a 98-label clinical vocabulary. With the same pre-trained retinal encoder and training budget, AMA attains a macro-AUPRC of 0.802 (95% CI: 0.788–0.823), compared with 0.545 (95% CI: 0.538–0.572) for the matched masked-supervision Baseline. These gains persist across four retinal backbones spanning masked image modeling, self-distillation, and vision–language pre-training, with improved per-label AUPRC estimates for 96 labels. The results support treating label observability as part of the learning problem when pooling medical datasets, rather than assuming that the pooled label matrix is complete.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.