GestaltCLIP: Concept-Enhanced Learning of Clinically Structured Facial Representation from Heterogeneous Phenotype Supervision
Abstract
Facial morphology provides important diagnostic cues for many rare genetic disorders, yet learning facial representations structured by clinically meaningful phenotypic concepts remains difficult because available datasets are small and supervision is highly fragmented. Rare disease facial representation learning is further limited by scarce image-level phenotype annotations, while richer but heterogeneous supervision exists across multiple levels, including Human Phenotype Ontology (HPO) concepts, patient phenotype sets, and disease level phenotype knowledge. We introduce **GestaltCLIP**, a concept-enhanced vision-language framework that integrates these heterogeneous phenotype signals to learn clinically structured facial representations without requiring complete image–phenotype pairing. GestaltCLIP unifies concept-, patient-, and disease-level phenotype supervision through ontology-aware HPO learning, phenotype-set modeling, direct image–phenotype grounding, and disease-mediated weak supervision. The resulting vision encoder operates with images alone at inference time. On GestaltMatcher Database, GestaltCLIP consistently improves disease recognition across four disease recognition metrics, yielding an average relative improvement of 15.58%, and achieves a 7.73-fold increase in mean average precision (mAP) for atomic image-to-HPO retrieval. This demonstrates substantially improved phenotype-concept grounding. Under strict external evaluation on RDFace, Top-1 accuracy improves from 8.93% to 15.18%. These results demonstrate that integrating heterogeneous phenotype supervision through explicit concept learning produces clinically structured, phenotype-grounded, and transferable facial representations for rare disease analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.