Holistic Diagnostic Semantics from Decomposed Imaging for Endoscopic Report Generation
Abstract
Endoscopic Report Generation (ERG) aims to automatically generate clinical reports from ear, nose, and throat (ENT) endoscopic images to assist clinical diagnosis. Unlike chest X-ray imaging, where a single image often provides relatively complete visual evidence, ENT endoscopy involves decomposed imaging, where multiple local views jointly support the overall diagnosis. Each view contains only partial anatomical and pathological information, while the clinical report summarizes findings from the entire examination. This discrepancy makes it challenging to derive holistic diagnostic semantics from distributed visual evidence and establish reliable correspondence between visual and textual representations. To address this challenge, we propose Holistic Semantic Learning for Endoscopic Report Generation (HSL-ERG), which learns holistic diagnostic semantics from distributed visual evidence under decomposed imaging. Specifically, the Holistic Semantic Representation Module (HSRM) aggregates distributed local visual evidence according to learned view contributions to construct a holistic examination-level semantic representation. Furthermore, the Reliable Correspondence Alignment Loss (RCAL) constructs adaptive bidirectional contrastive constraints between visual and textual representations to improve their correspondence reliability. Experiments on ENT-Endo, OCASD, and CTRG-Brain demonstrate the effectiveness of HSL-ERG. On ENT-Endo, HSL-ERG outperforms existing methods by 0.010 in BLEU-4 and 0.020 in clinical F1.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.