acceptodds
Under review as a conference paper at ICLR 2027

Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis

Abstract

Rare-disease diagnosis is inherently a long-tail reasoning (“Odyssey") problem: individual disorders are sparsely documented, clinical phenotypes are heterogeneous and incomplete, and relevant evidence is fragmented across disease ontologies, phenotype databases, gene annotations, and biomedical literature. As a result, language models often over-prioritize common diseases, overlook rare candidates, or produce plausible but invalid diagnoses. To address these challenges, we introduce RareDx, a unified large-language-model (LLM)-based agent for evaluating and improving rare-disease diagnostic reasoning by organizing heterogeneous medical knowledge into structured, verifiable evidence and distilling this knowledge into compact language models. RareDX has its own harness, where RareDx-Harness normalizes heterogeneous patient records into ranked differential-diagnosis tasks and systematically compares parametric inference, static retrieval augmentation, adaptive tool use, and structured phenotype–gene–disease reasoning over a shared medical knowledge infrastructure. Our analysis reveals that retrieval is not universally beneficial: static retrieval can introduce distracting or weakly relevant evidence, whereas adaptive retrieval and ontology-guided reasoning more consistently improve diagnostic ranking. Building on these findings, we develop a two-stage training pipeline combining Top-10 supervised fine-tuning with knowledge-graph–embedded reinforcement learning. The reward function projects model predictions into a structured medical knowledge space and jointly incorporates graded diagnostic relevance, disease-ontology proximity, biomedical semantic similarity, and disease–phenotype graph consistency, thereby transforming sparse exact-match feedback into dense and clinically meaningful supervision. Vocabulary constraints and contamination-aware penalties further suppress fabricated diagnoses and reward hacking. Under our benchmark study, our post-trained LLMs RareDx-9B and RareDx-28B surpass GPT-5.5 and other frontier models, demonstrating that structured knowledge integration and knowledge-aware post-training can enable compact models to achieve strong performance in long-tail rare-disease diagnostic reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.