acceptodds
Under review as a conference paper at ICLR 2027

Phenotype Driven Gene Prioritization and Fast Patient Retrieval with Rotational Knowledge Graph Embeddings

Abstract

The central challenge in variant interpretation for rare diseases is identifying the causal gene from a pool of candidates based on the patient's phenotypes. It is heavily dependent on subject domain expertise and is constrained to associations already documented in curated databases and the literature. In this study, we introduce a novel method - Patient Retrieval, Interpretation, and gene Scoring Model (PRISM), which is geometry aware, permutation invariant Set Transformer encoder that aligns variable size sets of patient phenotypes or knowledge graph (KG) entities, to candidate entities in rotational KG embedding (KGE) space. The KGE was built with the RotatE model over gene-phenotype and phenotype-phenotype relations from the Human Phenotype Ontology (HPO). Entity text descriptions were encoded with the pretrained biomedical language model, conferring inductive capability on PRISM. PRISM was trained on phenotypic profiles of approximately 22,000 rare disease patients, drawn from an in-house cohort of ∼30,000, to prioritize causal genes with a listwise multi-positive softmax loss function. On a held out cohort of more than 400 samples, integrated into the proprietary variant prioritization pipeline, PRISM recovered the causative pathogenic/likely pathogenic gene and variant within the Top-20 in 92.2% of cases and ranked the causative variant Top-1 in 57.8% of cases. Beyond the causal gene prioritization score, the same encoder embeds each patient's phenotypic profile in under half a second per inference, enabling retrieval of prior cases with similar phenotypes to aid variant interpretation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.