PrimeKG-CL: A Continual Graph Learning Benchmark on Evolving Biomedical Knowledge Graphs
Abstract
Biomedical knowledge graphs underwrite drug repurposing and clinical decision support, yet the upstream ontologies they depend on (GO, HPO, MONDO, CTD) update on independent cycles that add millions of edges and deprecate hundreds of thousands more between releases. Yet existing continual graph learning (CGL) has been studied almost exclusively on synthetic random splits of static, generic KGs, a regime that cannot reproduce the asynchronous, structured evolution real biomedical KGs undergo. To this end, we introduce PrimeKG-CL, a CGL benchmark built on PrimeKG (integrating 20+ biomedical databases; 129K+ nodes, 8.1M+ edges, 10 node types, 30 relation types) with a second temporal snapshot reconstructed from nine freely re-queried databases ( June 2021, July 2023; 5.76M edges added, 889K removed, 7.21M persistent), six entity-type-grouped tasks tracking real additions, multimodal node features (BiomedBERT text, Morgan fingerprints, R-GCN structural), and a per-task persistent/removed/added test stratification. On three tasks (biomedical relationship prediction, entity classification, KGQA), we evaluate six CL strategies across four KGE decoders, plus LKGE, an LLM-RAG agent, and CMKL. We find that decoder choice dominates the continual-learning strategy: link-prediction accuracy spans roughly across decoders (RotatE DistMult ComplEx TransE) versus up to 4 across CL strategies, and the two interact so that no single strategy is best across decoders (e.g., EWC markedly benefits ComplEx while Distillation degrades RotatE). Moreover, only DistMult exhibits a clear separation between persistent and deprecated knowledge ( higher MRR on persistent vs. removed triples), indicating that standard metrics conflate retention of still-valid facts with failure to forget outdated ones; this effect is absent under RotatE. In addition, multimodal fusion improves entity classification by over feature-matched continual baselines (Macro-F1 vs. , ), and a recent CKGE framework (IncDE) failed to scale to our 5.67M-triple base task across five attempts up to 350 GB RAM. Data, pipeline, baselines, and the stratified split will be released openly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.