RuralDoc: Evidence-Based Knowledge Graph Retrieval for Skin Disease Reasoning
Abstract
Structured external knowledge is often introduced to improve large language model (LLM) reasoning, yet poorly integrated knowledge can instead override correct model beliefs. We study this failure mode in RuralDoc, a multimodal skin-disease reasoning framework that integrates retrieval-augmented generation (RAG) with a hierarchical diagnostic knowledge graph (KG). An initial ranking-based design presented KG predictions directly to the LLM and degraded performance for most backbones we tested, with deference analysis showing that models frequently abandoned correct initial answers when the KG ranking was wrong. We therefore redesign the KG as an evidence provider rather than a diagnostic authority: the revised method uses IDF-weighted matching over diagnostic symptoms and visual features, withholds disease rankings and scores from the LLM, and presents retrieved facts as unordered reference evidence alongside dermatology context. Across 192 diagnostic cases and five language models, the combined redesign removes the degradation observed with the original KG integration and improves Top-1 accuracy from 70.3% to 78.6% for Llama-3.1-8B and from 88.1% to 93.8% for DeepSeek-V3.1 relative to model-only baselines. In cases where the KG's top prediction is wrong but the model's initial answer is correct, erroneous deference drops sharply, from 67% to 7% for Llama-3.1-8B, with similar reductions for the other four models. These results show that the way structured knowledge is retrieved and exposed to an LLM can materially affect downstream reasoning and motivate treating KG outputs as supporting evidence rather than authoritative predictions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.